Jaewon Kwon

dblp:362/8745 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0000-6352-0546ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Toward Scalable Gate-Level Parallelism on Trapped-Ion Processors with Racetrack Electrodes
abstract
A recent advancement in quantum computing shows a quantum advantage of certified randomness on the racetrack processor. This work investigates the execution efficiency of this architecture for general-purpose programs. We first explore the impact of increasing zones on runtime efficiency. Counterintuitively, our evaluations using variational programs reveal that expanding zones may degrade runtime performance under the existing scheduling policy. This degradation may be attributed to the increase in track length, which increases ion circulation overhead, offsetting the benefits of enhanced parallelism. To mitigate this, the proposed Plutarch exploits 3 strategies: (i) unitary decomposition and translation to maximize zone utilization, (ii) prioritizing the execution of nearby gates over ion circulation, and (iii) implementing shortcuts to provide the alternative path.
Enhyeok Jang, Hyungseok Kim 0003, Yongju Lee 0003, Jaewon Kwon, Yipeng Huang 0001, Won Woo Ro
HPCA4
2025 Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server Workloads
abstract
Modern CPUs suffer from the frontend bottleneck because the instruction footprint of server workloads exceeds the private cache capacity.Prior works have examined the CPU components or private cache to improve the instruction hit rate.The large footprint leads to significant cache misses not only in the core and faster-level cache but also in the last-level cache (LLC).We observe that even with an advanced branch predictor and instruction prefetching techniques, a considerable amount of instruction accesses descend to the LLC.However, state-of-the-art LLC designs with elaborate data management overlook handling the instruction misses that precede corresponding data accesses.Specifically, when an instruction requiring numerous data accesses is missed, the frontend of a CPU should wait for the instruction fetch, regardless of how much data are present in the LLC.To preserve hot instructions in the LLC, we propose Garibaldi, a novel pairwise instruction-data management scheme.Garibaldi tracks the hotness of instruction accesses by coupling it with that of data accesses and adopts management techniques.On the one hand, this scheme includes a selective protection mechanism that prevents the cache evictions of high-cost instruction cachelines.On the other hand, in the case of unprotected instruction line misses, Garibaldi conservatively issues prefetch requests of the paired data lines while handling those misses.In our experiments, we evaluate Garibaldi with 16 server workloads on a 40-core machine.We also implement Garibaldi on top of a modern LLC design, including Mockingjay.Garibaldi improves 13.2% and 6.1% of CPU performance on baseline LLC design and Mockingjay, respectively.
Jaewon Kwon, Yongju Lee 0003, Enhyeok Jang, Hongju Kal, Won Woo Ro
ISCA1
2025 COSMOS: An LLC Contention Slowdown Model for Heterogeneous Multi-Core Systems
abstract
Heterogeneous multi-core systems are increasingly adopted due to their advantages in area efficiency and energy savings. However, existing analytical models often overlook core heterogeneity, leading to lower performance prediction accuracy compared to homogeneous systems. In this paper, we show that even under identical last-level cache (LLC) contention conditions, heterogeneous cores experience different slowdowns. We categorize memory access time into internal and external components based on whether memory requests are served before reaching LLC and analyze how these two types affect application slowdowns. Furthermore, we examine how these components vary with core heterogeneity. Our analysis reveals that differences in cache hierarchies lead to distinct eviction patterns and variable external accesses, producing LLC miss rates that depend on LLC capacity. Additionally, core heterogeneity influences the execution times of both computation and internal memory accesses, which serve as correction factors that modulate the effect of LLC miss rate differences on application slowdown. Based on these insights, we propose COSMOS, an analytical model designed to accurately predict slowdowns caused by LLC contention in heterogeneous multi-core systems. COSMOS profiles the sensitivity of external accesses to LLC capacity, estimates LLC miss rates and average access latency, and aggregates the weighted contributions of all components. COSMOS achieves an average accuracy of 94.71% in performance prediction, significantly outperforming models that overlook internal resources, which achieve average accuracies of 82.76 % and 89.87 %, respectively.
Yongju Lee 0003, Jaewon Kwon, Cheolhwan Kim, Enhyeok Jang, Jiwon Lee 0001, Hyunwuk Lee, Won Woo Ro
ISPASS2
2025 HashScape: Leveraging Virtual Address Dynamics for Efficient Hashed Page Tables
abstract
The evolving memory landscape for larger capacity prompts alternative approaches due to scalability challenges in multi-level page tables, which require multiple serial memory accesses for address translation. Hashed Page Tables (HPTs) have gained attention for ideally facilitating a single memory access per translation. However, current HPTs increase minor page fault latency, thereby impeding its superiority over conventional multi-level page table design. This paper provides a comprehensive analysis of HPTs regarding minor page fault latency concerning memory management subsystems. In particular, we demonstrate how feasibility issues in memory management with HPTs can escalate minor page fault latency. We observe that different page types in HPTs (anon pages and page caches) exhibit distinct behaviors on the occurrence of minor page faults, indicating a significant correlation between page types and minor page faults. To address these challenges, we proposeHashScape, a scheme that harmonizes with memory management using tailored HPTs per segment and size-tailored allocation via Virtual Memory Areas. Our evaluation demonstrates that HashScape significantly improves the insertion latency, with average, 95th, and 99thpercentiles improving by 1.8$\boldsymbol{\times}$, 1.9$\boldsymbol{\times}$, and 2.2$\boldsymbol{\times}$, respectively, resulting in an overall 10% reduction in minor page fault latency compared to a state-of-the-art HPT design.
Won Hur, Jiwon Lee 0001, Jaewon Kwon, Minjae Kim 0010, Won Woo Ro
IEEE Trans. Computers3
2024 Recompiling QAOA Circuits on Various Rotational Directions
abstract
The quantum approximate optimization algorithm (QAOA) is introduced to efficiently solve combinatorial optimization problems. Despite the promise of QAOA, the cost of executing QAOA circuits at scale for quantum advantage may still be excessive for the near-future quantum device. We observe the increasing overhead of QAOA circuit execution in the native gate translation. To execute QAOA circuits on a real quantum computing device, Hamiltonians composed of predefined specific rotations (e.g., ZZ and X) should be decomposed into finite native gates. By adopting rotational combinations that utilize native gates more directly than the standard QAOA circuit model, the execution cost on real quantum devices can be reduced. In this study, we propose Racoon (Rotational Space Virtualization for QAOA Ansatz), an algorithm-hardware co-design approach that revisits the synthesis conditions of QAOA circuits and selects alternative candidates with different rotational combinations. Our analysis of six commercial quantum processors demonstrates that applying Racoon to QAOA circuits for the 4-node Sherrington-Kirkpatrick model reduces the number of native gates by an average of 23% and up to 79%. Consequently, using Racoon results in 43% fewer training epochs, 41% lower training energy consumption, and a 6% improvement in inference on average compared to standard QAOA. Racoon consistently reduces circuit depth as the number of qubits and layers increases, achieving 123 × more circuit depth reduction compared to the recently proposed Depth First Search (DFS)-based method. Furthermore, we confirm that Racoon’s method can be extended to State-of-The-Art QAOAs with modified ansätze and to the variational quantum eigensolver (VQE).
Enhyeok Jang, Dongho Ha, Seungwoo Choi 0001, Youngmin Kim 0005, Jaewon Kwon, Yongju Lee 0003, Sungwoo Ahn, Hyungseok Kim 0003, Won Woo Ro
PACT5
2024 Geneva: A Dynamic Confluence of Speculative Execution and In-Order Commitment Windows
abstract
Modern out-of-order microprocessors are increasingly expanding resources such as reorder buffer (ROB) and instruction queue (IQ) for memory-level parallelism (MLP). While this expansion effectively addresses the memory wall challenge, it also incurs notable cost and energy trade-offs. To tackle this, we propose Geneva, a microarchitecture that improves performance and saves energy. Geneva reallocates a portion of an ROB to serve as a dynamic queue (DQ), used as an ROB, IQ, or both depending on operational needs. Geneva saves energy by 15.6% and improves performance by 2.6% compared to the conventional out-of-order core.
Yanghee Lee, Jiwon Lee 0001, Jaewon Kwon, Yongju Lee 0003, Won Woo Ro
DAC3
2023 McCore: A Holistic Management of High-Performance Heterogeneous Multicores
abstract
Heterogeneous multicore systems have emerged as a promising approach to scale performance in high-end desktops within limited power and die size constraints. Despite their advantages, these systems face three major challenges: memory bandwidth limitation, shared cache contention, and heterogeneity. Small cores in these systems tend to occupy a significant portion of shared LLC and memory bandwidth, despite their lower computational capabilities, leading to performance degradation of up to 18% in memory-intensive workloads. Therefore, it is crucial to address these challenges holistically, considering shared resources and core heterogeneity while managing shared cache and bandwidth.
Jaewon Kwon, Yongju Lee 0003, Hongju Kal, Minjae Kim 0010, Youngsok Kim, Won Woo Ro
MICRO1