EDBT 2026 Demo / reviewers in the wild / expert
Jiwon Lee 0001
dblp:76/9606-1
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2025
0009-0009-6529-5333ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Marching Page Walks: Batching and Concurrent Page Table Walks for Enhancing GPU ThroughputabstractVirtual memory, with the support of address translation hardware, is a key technique in expanding programmability and memory management in GPUs. However, the nature of the GPU execution model heavily pressures its translation hardware, particularly due to a discrepancy in the behavior of page table walkers and thousands of concurrently running threads. In GPU workloads, multiple threads simultaneously access a number of pages necessitating a substantial number of translations whereas each walker handles only a single walk request at a time. Such a limitation significantly increases the queueing latency of walk requests, which we observe as a major bottleneck for servicing page table walks in GPUs. To tackle this challenge, we investigate a design of page walkers that facilitates multiple walk requests to be handled together in batches. Then, we make the following observations: 1) allowing a page walker to issue beyond a single memory request significantly improves the throughput of walkers, and 2) GPU applications tend to concurrently access pages in wide address ranges. By leveraging the above implications, we propose Marching Page Walks (MPW) that effectively mitigate the contention in GPU page table walkers. MPW scans pending walk requests to identify ones that can be grouped together. Then, MPW batches these requests and concurrently handles them by issuing multiple memory instructions. Experiments show that MPW reduces the queueing latency of page walks by 86.7% and improves GPU performance by 55.6% over the baseline design. Jiwon Lee 0001, Gun Ko, Myung Kuk Yoon, Ipoom Jeong, Yunho Oh, Won Woo Ro |
HPCA | 1 |
| 2025 | Heliostat: Harnessing Ray Tracing Accelerators for Page Table WalksabstractThis paper introduces Heliostat, which enhances page translation bandwidth on GPUs by harnessing underutilized ray tracing accelerators (RTAs).While most existing studies focused on better utilizing the provided translation bandwidth, this paper introduces a new opportunity to fundamentally increase the translation bandwidth.Instead of overprovisioning the GPU memory management unit (GMMU), Heliostat repurposes the existing RTAs by leveraging the operational similarities between ray tracing and page table walks.Unlike earlier studies that utilized RTAs for certain workloads, Heliostat democratizes RTA for supporting any workloads by improving virtual memory performance.Heliostat+ optimizes Heliostat by handling predicted future address translations proactively.Heliostat outperforms baseline and two state-of-the-arts by 1.93×, 1.92×, and 1.66×.Heliostat+ further speeds up Heliostat by 1.23×.Compared to an overprovisioned comparable solution, Heliostat occupies only 1.53% of the area and consumes 5.8% of the power. Yuke Li 0003, Jiwon Lee 0001, Won Woo Ro, Hyeran Jeon |
ISCA | 3 |
| 2025 | COSMOS: An LLC Contention Slowdown Model for Heterogeneous Multi-Core SystemsabstractHeterogeneous multi-core systems are increasingly adopted due to their advantages in area efficiency and energy savings. However, existing analytical models often overlook core heterogeneity, leading to lower performance prediction accuracy compared to homogeneous systems. In this paper, we show that even under identical last-level cache (LLC) contention conditions, heterogeneous cores experience different slowdowns. We categorize memory access time into internal and external components based on whether memory requests are served before reaching LLC and analyze how these two types affect application slowdowns. Furthermore, we examine how these components vary with core heterogeneity. Our analysis reveals that differences in cache hierarchies lead to distinct eviction patterns and variable external accesses, producing LLC miss rates that depend on LLC capacity. Additionally, core heterogeneity influences the execution times of both computation and internal memory accesses, which serve as correction factors that modulate the effect of LLC miss rate differences on application slowdown. Based on these insights, we propose COSMOS, an analytical model designed to accurately predict slowdowns caused by LLC contention in heterogeneous multi-core systems. COSMOS profiles the sensitivity of external accesses to LLC capacity, estimates LLC miss rates and average access latency, and aggregates the weighted contributions of all components. COSMOS achieves an average accuracy of 94.71% in performance prediction, significantly outperforming models that overlook internal resources, which achieve average accuracies of 82.76 % and 89.87 %, respectively. Yongju Lee 0003, Jaewon Kwon, Cheolhwan Kim, Enhyeok Jang, Jiwon Lee 0001, Hyunwuk Lee, Won Woo Ro |
ISPASS | 5 |
| 2025 | LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR Compression
Yeonan Ha, Hanna Cha, Jiwon Lee 0001, Joonsung Kim 0001, Won Woo Ro, Youngsok Kim |
MICRO | 4 |
| 2025 | REC: Enhancing fine-grained cache coherence protocol in multi-GPU systems
Gun Ko, Jiwon Lee 0001, Hongju Kal, Hyunwuk Lee, Won Woo Ro |
J. Syst. Archit. | 2 |
| 2025 | HashScape: Leveraging Virtual Address Dynamics for Efficient Hashed Page TablesabstractThe evolving memory landscape for larger capacity prompts alternative approaches due to scalability challenges in multi-level page tables, which require multiple serial memory accesses for address translation. Hashed Page Tables (HPTs) have gained attention for ideally facilitating a single memory access per translation. However, current HPTs increase minor page fault latency, thereby impeding its superiority over conventional multi-level page table design. This paper provides a comprehensive analysis of HPTs regarding minor page fault latency concerning memory management subsystems. In particular, we demonstrate how feasibility issues in memory management with HPTs can escalate minor page fault latency. We observe that different page types in HPTs (anon pages and page caches) exhibit distinct behaviors on the occurrence of minor page faults, indicating a significant correlation between page types and minor page faults. To address these challenges, we proposeHashScape, a scheme that harmonizes with memory management using tailored HPTs per segment and size-tailored allocation via Virtual Memory Areas. Our evaluation demonstrates that HashScape significantly improves the insertion latency, with average, 95th, and 99thpercentiles improving by 1.8$\boldsymbol{\times}$, 1.9$\boldsymbol{\times}$, and 2.2$\boldsymbol{\times}$, respectively, resulting in an overall 10% reduction in minor page fault latency compared to a state-of-the-art HPT design. Won Hur, Jiwon Lee 0001, Jaewon Kwon, Minjae Kim 0010, Won Woo Ro |
IEEE Trans. Computers | 2 |
| 2024 | Geneva: A Dynamic Confluence of Speculative Execution and In-Order Commitment WindowsabstractModern out-of-order microprocessors are increasingly expanding resources such as reorder buffer (ROB) and instruction queue (IQ) for memory-level parallelism (MLP). While this expansion effectively addresses the memory wall challenge, it also incurs notable cost and energy trade-offs. To tackle this, we propose Geneva, a microarchitecture that improves performance and saves energy. Geneva reallocates a portion of an ROB to serve as a dynamic queue (DQ), used as an ROB, IQ, or both depending on operational needs. Geneva saves energy by 15.6% and improves performance by 2.6% compared to the conventional out-of-order core. Yanghee Lee, Jiwon Lee 0001, Jaewon Kwon, Yongju Lee 0003, Won Woo Ro |
DAC | 2 |
| 2023 | SnakeByte: A TLB Design with Adaptive and Recursive Page Merging in GPUsabstractThis paper presents an address translation scheme in GPUs named SnakeByte that can dynamically manage variable-sized pages and maximize TLB reach by recursively merging contiguous pages. Memory virtualization has become an integral part of GPUs to enhance programmability and memory management efficiency. However, conventional memory virtualization methods using multi-level page tables and caching them in TLBs are insufficient to provide GPUs with enough address translation coverage for the massive volume of data. SnakeByte implements a hardware-based address translation mechanism that recursively merges contiguous pages into larger page groups and effectively extends TLB coverage. SnakeByte allows multiple equal-sized pages coalescing into a page table entry (PTE). It records the validity of pages to be merged using a bit vector, and few bits are annexed to indicate the size of merged pages. If all pages covered by the PTE are allocated with contiguity, the PTE is promoted to be further coalesced into a larger page group. The recursive coalescence of contiguous pages enables SnakeByte to handle variable-sized page groups with the exponentially increasing TLB reach. Associated with a contiguity-aware memory allocator, SnakeByte can consolidate vastly contiguous address spaces into a few TLB entries. Consequently, it significantly reduces TLB misses for large working sets in GPUs and achieves substantial performance improvements. Experiment results show that SnakeByte decreases the number of page table walks by 6.5x and enhances the GPU performance by 2.0x on average over the conventional paging scheme. Jiwon Lee 0001, Ju Min Lee, Yunho Oh, William J. Song, Won Woo Ro |
HPCA | 1 |
| 2023 | Early-Adaptor: An Adaptive Framework forProactive UVM Memory ManagementabstractUnified Virtual Memory (UVM) relieves programmers of the burden of memory management between CPU and GPUs. However, the use of UVM can lead to performance degradation due to its on-demand page migration scheme, especially under memory oversubscription. In this research, we conduct various analyses on real hardware, NVIDIA RTX 3090, to examine such performance degradation with an NVIDIA opensource GPU driver. Our analysis shows that the effectiveness of prefetching highly correlates with the relative number of page faults on a group of contiguous pages, which NVIDIA refers to as a Virtual Address Block (VABlock) spanning across a 2MB virtual address range. Also, the risk of page thrashing is determined by the total number of VABlocks that consistently generate page faults during kernel execution. Hence, the performance impact of the prefetch threshold varies across different workloads. These observations indicate that an adaptive prefetching scheme can resolve the performance bottleneck of memory oversubscription. To this end, we propose the Early-Adaptor (EA) framework, which automatically controls the prefetching aggressiveness based on the page fault history. During runtime, the EA framework monitors patterns of page faults in per-VABlock and in a global scope. After analyzing page fault generation rates and the possibility of page thrashing, the EA framework dynamically controls the prefetching aggressiveness by changing the prefetch threshold. The EA framework requires only minor changes to GPU drivers and needs no changes to the GPU hardware. Experiments on real hardware show that when GPU memory is oversubscribed, the EA framework achieves an average speedup of 1. 74x over the conventional GPU prefetcher. Seokjin Go, Hyunwuk Lee, Junsung Kim 0002, Jiwon Lee 0001, Myung Kuk Yoon, Won Woo Ro |
ISPASS | 4 |
| 2022 | Reconstructing Out-of-Order Issue QueueabstractOut-of-order cores provide high performance at the cost of energy efficiency. Dynamic scheduling is one of the major contributors to this: generating highly optimized issue schedules considering both data dependences and underlying execution resources, but relying heavily on complex wakeup and select operations of an out-of-order issue queue (IQ). For decades, researchers have proposed several complexity-effective dynamic scheduling schemes by leveraging the energy efficiency of an in-order IQ. However, they are either costly or not capable of delivering sufficient performance to substitute for a conventional wide-issue out-of-order IQ. In this work, we revisit two previous designs: one classical dependence-based design and the other state-of-the-art readiness-based design. We observe that they are complementary to each other, and thus their synergistic integration has the potential to be a good alternative to an out-of-order IQ. We first combine these two designs, and further analyze the main architectural bottlenecks that incur the underutilization of aggregate issue capability, thereby limiting the exploitation of instruction-level and memory-level parallelisms: 1) memory dependences not exposed by the register-based dependence analysis and 2) wide and shallow nature of dynamic dependence chains due to the long-latency memory accesses. To this end, we propose Ballerino, a novel microarchitecture that performs balanced and cache-miss-tolerable dynamic scheduling via a complementary combination of cascaded and clustered in-order IQs. Ballerino is built upon three key functionalities: 1) speculatively filtering out ready-at-dispatch instructions, 2) eliminating wasteful wakeup operations via a simple steering technique leveraging the awareness of memory dependences, and 3) reacting to program phase changes by allowing different load-dependent chains to share a single IQ while guaranteeing their out-of-order issue. The net effect is minimal scheduling energy consumption per instruction while providing comparable scheduling performance to a fully out-of-order IQ. In our analysis, Ballerino achieves comparable performance to an 8-wide out-of-order core by using twelve in-order IQs, improving core-wide energy efficiency by 20%. Ipoom Jeong, Jiwon Lee 0001, Myung Kuk Yoon, Won Woo Ro |
MICRO | 2 |