EDBT 2026 Demo / reviewers in the wild / expert
Hongju Kal
dblp:297/1496
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0001-8443-1734ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server WorkloadsabstractModern CPUs suffer from the frontend bottleneck because the instruction footprint of server workloads exceeds the private cache capacity.Prior works have examined the CPU components or private cache to improve the instruction hit rate.The large footprint leads to significant cache misses not only in the core and faster-level cache but also in the last-level cache (LLC).We observe that even with an advanced branch predictor and instruction prefetching techniques, a considerable amount of instruction accesses descend to the LLC.However, state-of-the-art LLC designs with elaborate data management overlook handling the instruction misses that precede corresponding data accesses.Specifically, when an instruction requiring numerous data accesses is missed, the frontend of a CPU should wait for the instruction fetch, regardless of how much data are present in the LLC.To preserve hot instructions in the LLC, we propose Garibaldi, a novel pairwise instruction-data management scheme.Garibaldi tracks the hotness of instruction accesses by coupling it with that of data accesses and adopts management techniques.On the one hand, this scheme includes a selective protection mechanism that prevents the cache evictions of high-cost instruction cachelines.On the other hand, in the case of unprotected instruction line misses, Garibaldi conservatively issues prefetch requests of the paired data lines while handling those misses.In our experiments, we evaluate Garibaldi with 16 server workloads on a 40-core machine.We also implement Garibaldi on top of a modern LLC design, including Mockingjay.Garibaldi improves 13.2% and 6.1% of CPU performance on baseline LLC design and Mockingjay, respectively. Jaewon Kwon, Yongju Lee 0003, Enhyeok Jang, Hongju Kal, Won Woo Ro |
ISCA | 5 |
| 2025 | REC: Enhancing fine-grained cache coherence protocol in multi-GPU systems
Gun Ko, Jiwon Lee 0001, Hongju Kal, Hyunwuk Lee, Won Woo Ro |
J. Syst. Archit. | 3 |
| 2023 | AESPA: Asynchronous Execution Scheme to Exploit Bank-Level Parallelism of Processing-in-MemoryabstractThis paper presents an asynchronous execution scheme to leverage the bank-level parallelism of near-bank processing-in-memory (PIM). We observe that performing memory operations underutilizes the parallelism of PIM computation because near-bank PIMs are designated to operate all banks synchronously. The all-bank computation can be delayed when one of the banks performs the basic memory commands, such as read/write requests and activation/precharge operations. We aim to mitigate the throughput degradation and especially focus on execution delay caused by activation/precharge operations. For all-bank execution accessing the same row of all banks, a large number of activation/precharge operations inevitably occur. Considering the timing parameter limiting the rate of row-open operations (tFAW), the throughput might decrease even further. To resolve this activation/precharge overhead, we propose AESPA, a new parallel execution scheme that operates banks asynchronously. AESPA is different from the previous synchronous execution in that (1) the compute command of AESPA targets a single bank, and (2) each processing unit computes data stored in multiple DRAM columns. By doing so, while one bank computes multiple DRAM columns, the memory controller issues activation/precharge or PIM compute commands to other banks. Thus, AESPA hides the activation latency of PIM computation and fully utilizes the aggregated bandwidth of the banks. For this, we modify hardware and software to support vector and matrix computation of previous near-bank PIM architectures. In particular, we change the matrix-vector multiplication based on an inner product to fit it on AESPA PIM. Previous matrix-vector multiplication requires data broadcasting and simultaneous computation across all processing units. By changing the matrix-vector multiplication method, AESPA PIM can transfer data to respective processing units and start computation asynchronously. As a result, the near-bank PIMs adopting AESPA achieve 33.5% and 59.5% speedup compared to two different state-of-the-art PIMs. Hongju Kal, Chanyoung Yoo, Won Woo Ro |
MICRO | 1 |
| 2023 | McCore: A Holistic Management of High-Performance Heterogeneous MulticoresabstractHeterogeneous multicore systems have emerged as a promising approach to scale performance in high-end desktops within limited power and die size constraints. Despite their advantages, these systems face three major challenges: memory bandwidth limitation, shared cache contention, and heterogeneity. Small cores in these systems tend to occupy a significant portion of shared LLC and memory bandwidth, despite their lower computational capabilities, leading to performance degradation of up to 18% in memory-intensive workloads. Therefore, it is crucial to address these challenges holistically, considering shared resources and core heterogeneity while managing shared cache and bandwidth. Jaewon Kwon, Yongju Lee 0003, Hongju Kal, Minjae Kim 0010, Youngsok Kim, Won Woo Ro |
MICRO | 3 |
| 2023 | A convertible neural processor supporting adaptive quantization for real-time neural networks
Hongju Kal, Hyoseong Choi, Ipoom Jeong, Joon-Sung Yang, Won Woo Ro |
J. Syst. Archit. | 1 |
| 2021 | SPACE: Locality-Aware Processing in Heterogeneous Memory for Personalized RecommendationsabstractPersonalized recommendation systems have become a major AI application in modern data centers. The main challenges in processing personalized recommendation inferences are the large memory footprint and high bandwidth requirement of embedding layers. To overcome the capacity limit and bandwidth congestion of on-chip memory, near memory processing (NMP) can be a promising solution. Recent work on accelerating personalized recommendations proposes a DIMMbased NMP design to solve the bandwidth problem and increases memory capacity. The performance of NMP is determined by the internal bandwidth and the prior DIMM-based approach utilizes more DIMMs to achieve higher operation throughput. However, extending the number of DIMMs could eventually lead to significant power consumption due to inefficient scaling. We propose SPACE, a novel heterogeneous memory architecture, which is efficient in terms of performance and energy. SPACE exploits a compute-capable 3D-stacked DRAM with DIMMs for personalized recommendations. Prior to designing the proposed system, we give a quantitative analysis of the user/item interactions and define the two localities: gather locality and reduction locality. In gather operations, we find only a small proportion of items are highly-accessed by users, and we call this gather locality. Also, we define reduction locality as the reusability of the gathered items in reduction operations. Based on the gather locality, SPACE allocates highly-accessed embedding items to the 3D-stacked DRAM to achieve the maximum bandwidth. Subsequently, by exploiting reduction locality, we utilize the remaining space of the 3D-stacked DRAM to store and reuse repeated partial sums, thereby minimizing the required number of element-wise reduction operations. As a result, the evaluation shows that SPACE achieves 3.2× performance improvement and 56% energy saving over the previous DIMM-based NMPs leveraging 3D-stacked DRAM with a 1/8 size of DIMMs. Also, compared to the state-of-the-art DRAM cache designs with the same NMP configuration, SPACE achieves an average 32.7% of performance improvement. Hongju Kal, Seokmin Lee, Gun Ko, Won Woo Ro |
ISCA | 1 |