Enhyeok Jang

dblp:357/2981 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0000-7034-6793ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 first-author · 10 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Toward Scalable Gate-Level Parallelism on Trapped-Ion Processors with Racetrack Electrodes
abstract
A recent advancement in quantum computing shows a quantum advantage of certified randomness on the racetrack processor. This work investigates the execution efficiency of this architecture for general-purpose programs. We first explore the impact of increasing zones on runtime efficiency. Counterintuitively, our evaluations using variational programs reveal that expanding zones may degrade runtime performance under the existing scheduling policy. This degradation may be attributed to the increase in track length, which increases ion circulation overhead, offsetting the benefits of enhanced parallelism. To mitigate this, the proposed Plutarch exploits 3 strategies: (i) unitary decomposition and translation to maximize zone utilization, (ii) prioritizing the execution of nearby gates over ion circulation, and (iii) implementing shortcuts to provide the alternative path.
Enhyeok Jang, Hyungseok Kim 0003, Yongju Lee 0003, Jaewon Kwon, Yipeng Huang 0001, Won Woo Ro
HPCA1
2026 Leveraging Phase Polynomials for Quantum Circuit Optimization
abstract
Quantum circuits on resource-limited hardware require optimizing regions dominated by $\{\mathrm{CNOT}, R_z\}$, which account for a large fraction of operations and often dominate execution cost. This optimization can be challenging because phase-polynomial blocks are fragmented by basis-changing gates such as $H$, and optimizing phase parities alone may increase the cost of downstream basis transformations. Existing phase-polynomial approaches are limited to single-block or phase-only optimization, while subcircuit rewriting approaches are local and scale poorly beyond small rewrite windows. We introduce \emph{PhasePoly}, a compiler optimization pass that jointly optimizes phase-parity and output-parity networks and employs a cross-block intermediate representation to reuse parities across phase-polynomial block barriers. This approach is effective because its unified parity-matrix representation exposes long-range $\{\mathrm{CNOT}, R_z\}$ structure that local rewriting and single-block methods cannot capture. \emph{PhasePoly} reduces total gate count by up to 50.00\% (34.70\% on average) and CNOT count by up to 48.57\% (26.83\% on average), while scaling to large circuits and improving both fault-tolerant compilation and near-term hardware execution. \emph{PhasePoly} is available at https://github.com/ruadapt/PhasePoly.
Zihan Chen 0005, Henry Chen, Yuwei Jin, Enhyeok Jang, Mingkuan Xu, Vannessa Chan, Won Woo Ro, Eddy Z. Zhang
ISCA4
2025 PIMutation: Exploring the Potential of Real PIM Architecture for Quantum Circuit Simulation
abstract
Quantum circuit simulations are essential for the verification of quantum algorithms on behalf of real quantum devices. However, the memory requirements for such simulations grow exponentially with the number of qubits involved in quantum programs. Moreover, a substantial number of computations in quantum circuit simulations cause low locality data accesses, as they require extensive computations across the entire table of the full state vector. These characteristics lead to significant latency and energy overheads during data transfers between the CPU and main memory. Processing-in-Memory (PIM), which integrates computational logic near DRAM banks, could present a promising solution to address these challenges.
Dongin Lee, Enhyeok Jang, Seungwoo Choi 0001, Junwoong An, Cheolhwan Kim, Won Woo Ro
ASP-DAC2
2025 Qubit Movement-Optimized Program Generation on Zoned Neutral Atom Processors
abstract
A zoned neutral atom architecture achieves exceptional fidelity by segregating the execution spaces of 1- and 2-qubit gates, being a promising candidate for high-accuracy quantum systems. Unfortunately, na'ively applying programs designed for static qubit topologies to zoned architectures may result in most execution time being consumed by intra-zone travels of atoms. To address this, we introduce Mantra (Minimizing trAp movemeNts for aTom aRray Architectures), which rewrites quantum programs to reduce the interleaving of single- and two-qubit gates. Mantra incorporates three strategies: (i) a fountain-shaped controlled-Z (CZ) chain, (ii) ZZ-interaction protocol without a 1-qubit gate, and (iii) preemptive gate scheduling. Mantra reduces inter-zone movements by 68%, physical gate counts by 35%, and improves circuit fidelities by 17% compared to the standard executions.
Enhyeok Jang, Youngmin Kim 0005, Hyungseok Kim 0003, Seungwoo Choi 0001, Yipeng Huang 0001, Won Woo Ro
CGO1
2025 QR-Map: A Map-Based Approach to Quantum Circuit Abstraction for Qubit Reuse Optimization
abstract
Recent advances in quantum computing introduce the ability to reuse qubits through mid-circuit measurements, thereby enhancing the efficiency of quantum devices with limited computational resources.However, identifying optimal reuse opportunities in quantum circuits remains challenging due to the intricate dependencies between quantum gates.Existing frameworks address this by either directly searching for reuse opportunities or converting circuits into directed acyclic graphs (DAGs).Unfortunately, these frameworks may require exponential search complexity or may not always ensure optimal results due to their non-deterministic property.To overcome these challenges, we propose QR-Map (Qubit Reuse Map), a map-based framework that abstracts computational dependencies for efficient qubit reuse.By extracting and aligning two-qubit gates, QR-Map facilitates dependency detection and ensures qubit savings without incurring excessive idle time.This approach achieves an optimal balance between gate serialization depth and crosstalk reduction.Evaluations with various quantum circuit benchmarks demonstrate that quantum circuits optimized with QR-Map achieve average reductions of 20% in qubit usage, 25% in circuit depth, and 22% in SWAP insertions compared to those optimized with the state-of-the-art framework.
Hyungseok Kim 0003, Enhyeok Jang, Seungwoo Choi 0001, Youngmin Kim 0005, Won Woo Ro
ISCA2
2025 Garibaldi: A Pairwise Instruction-Data Management for Enhancing Shared Last-Level Cache Performance in Server Workloads
abstract
Modern CPUs suffer from the frontend bottleneck because the instruction footprint of server workloads exceeds the private cache capacity.Prior works have examined the CPU components or private cache to improve the instruction hit rate.The large footprint leads to significant cache misses not only in the core and faster-level cache but also in the last-level cache (LLC).We observe that even with an advanced branch predictor and instruction prefetching techniques, a considerable amount of instruction accesses descend to the LLC.However, state-of-the-art LLC designs with elaborate data management overlook handling the instruction misses that precede corresponding data accesses.Specifically, when an instruction requiring numerous data accesses is missed, the frontend of a CPU should wait for the instruction fetch, regardless of how much data are present in the LLC.To preserve hot instructions in the LLC, we propose Garibaldi, a novel pairwise instruction-data management scheme.Garibaldi tracks the hotness of instruction accesses by coupling it with that of data accesses and adopts management techniques.On the one hand, this scheme includes a selective protection mechanism that prevents the cache evictions of high-cost instruction cachelines.On the other hand, in the case of unprotected instruction line misses, Garibaldi conservatively issues prefetch requests of the paired data lines while handling those misses.In our experiments, we evaluate Garibaldi with 16 server workloads on a 40-core machine.We also implement Garibaldi on top of a modern LLC design, including Mockingjay.Garibaldi improves 13.2% and 6.1% of CPU performance on baseline LLC design and Mockingjay, respectively.
Jaewon Kwon, Yongju Lee 0003, Enhyeok Jang, Hongju Kal, Won Woo Ro
ISCA4
2025 COSMOS: An LLC Contention Slowdown Model for Heterogeneous Multi-Core Systems
abstract
Heterogeneous multi-core systems are increasingly adopted due to their advantages in area efficiency and energy savings. However, existing analytical models often overlook core heterogeneity, leading to lower performance prediction accuracy compared to homogeneous systems. In this paper, we show that even under identical last-level cache (LLC) contention conditions, heterogeneous cores experience different slowdowns. We categorize memory access time into internal and external components based on whether memory requests are served before reaching LLC and analyze how these two types affect application slowdowns. Furthermore, we examine how these components vary with core heterogeneity. Our analysis reveals that differences in cache hierarchies lead to distinct eviction patterns and variable external accesses, producing LLC miss rates that depend on LLC capacity. Additionally, core heterogeneity influences the execution times of both computation and internal memory accesses, which serve as correction factors that modulate the effect of LLC miss rate differences on application slowdown. Based on these insights, we propose COSMOS, an analytical model designed to accurately predict slowdowns caused by LLC contention in heterogeneous multi-core systems. COSMOS profiles the sensitivity of external accesses to LLC capacity, estimates LLC miss rates and average access latency, and aggregates the weighted contributions of all components. COSMOS achieves an average accuracy of 94.71% in performance prediction, significantly outperforming models that overlook internal resources, which achieve average accuracies of 82.76 % and 89.87 %, respectively.
Yongju Lee 0003, Jaewon Kwon, Cheolhwan Kim, Enhyeok Jang, Jiwon Lee 0001, Hyunwuk Lee, Won Woo Ro
ISPASS4
2024 Recompiling QAOA Circuits on Various Rotational Directions
abstract
The quantum approximate optimization algorithm (QAOA) is introduced to efficiently solve combinatorial optimization problems. Despite the promise of QAOA, the cost of executing QAOA circuits at scale for quantum advantage may still be excessive for the near-future quantum device. We observe the increasing overhead of QAOA circuit execution in the native gate translation. To execute QAOA circuits on a real quantum computing device, Hamiltonians composed of predefined specific rotations (e.g., ZZ and X) should be decomposed into finite native gates. By adopting rotational combinations that utilize native gates more directly than the standard QAOA circuit model, the execution cost on real quantum devices can be reduced. In this study, we propose Racoon (Rotational Space Virtualization for QAOA Ansatz), an algorithm-hardware co-design approach that revisits the synthesis conditions of QAOA circuits and selects alternative candidates with different rotational combinations. Our analysis of six commercial quantum processors demonstrates that applying Racoon to QAOA circuits for the 4-node Sherrington-Kirkpatrick model reduces the number of native gates by an average of 23% and up to 79%. Consequently, using Racoon results in 43% fewer training epochs, 41% lower training energy consumption, and a 6% improvement in inference on average compared to standard QAOA. Racoon consistently reduces circuit depth as the number of qubits and layers increases, achieving 123 × more circuit depth reduction compared to the recently proposed Depth First Search (DFS)-based method. Furthermore, we confirm that Racoon’s method can be extended to State-of-The-Art QAOAs with modified ansätze and to the variational quantum eigensolver (VQE).
Enhyeok Jang, Dongho Ha, Seungwoo Choi 0001, Youngmin Kim 0005, Jaewon Kwon, Yongju Lee 0003, Sungwoo Ahn, Hyungseok Kim 0003, Won Woo Ro
PACT1
2024 Barber: Balancing Thermal Relaxation Deviations of NISQ Programs by Exploiting Bit-Inverted Circuits
abstract
One of the predominant causes of program distortion in the real quantum computing system may be attributed to the probability deviation caused by thermal relaxation. We introduce Barber (Balancing reAdout Results using Bit-invErted ciRcuits), a method designed to counteract the asymmetric thermal relaxation deviation and improve the reliability of near-term quantum programs. Barber collaborates with a bit-inverted quantum circuit, where the excited quantum state of qubits is assigned to the |0〉 and the unexcited state to the |1〉. In doing so, bit-inverted quantum circuits can experience thermal relaxation in the opposite direction compared to standard quantum circuits. Barber can effectively suppress the thermal relaxation deviation in program's readout results by selectively merging distributions from the standard and bit-inverted circuits.
Enhyeok Jang, Seungwoo Choi 0001, Youngmin Kim 0005, Jeewoo Seo, Won Woo Ro
ICCAD1
2024 MOSQ: Accelerating Classical Simulation of UCCSD Ansatz Circuits using Merged Operation
abstract
The Variational Quantum Eigensolver (VQE) is considered one of the most effective algorithms for near-term quantum processors due to its potential to produce meaningful results and its relatively small number of required qubits. However, the Unitary Coupled Cluster Singles and Doubles (UCCSD) circuit, used as the ansatz circuit for VQE, requires an excessive number of gate operations. This consequently causes long simulation delay when we simulate any VQE algorithm on classical computers. In order to enhance this long simulation delay of VQE, we develop and demonstrate that each Pauli string composing the UCCSD circuit can be merged into a single operation and executed efficiently in a classical simulator. We propose MOSQ, Merged Operation in Sub Quantum circuits for Pauli strings, which operates in a coupled manner in the circuit compiler stage and the execution stage to utilize merged operations. MOSQ passes the Pauli string information from the compiler stage to the execution stage, where each Pauli string is computed in the execution stage as a merged operation that functions similarly to a l-qubit gate operation except for the memory access pattern. MOSQ shows$12.2\times$and$8.67\times$speedup in UCCSD simulation time and total VQE execution time, respectively, compared to the baseline qiskit-aer simulator. Additionally, it is$4.88\times$and$3.11\times$faster than qiskit-aer simulator with fusion optimization enabled.
Seungwoo Choi 0001, Enhyeok Jang, Youngmin Kim 0005, Won Woo Ro
ICCD2
2023 Quixote: Improving Fidelity of Quantum Program by Independent Execution of Controlled Gates
abstract
NISQ (noisy intermediate-scale quantum) computers are vulnerable to errors, which limit the size of verifiable quantum circuits. For large quantum circuits, it is more difficult to obtain reliable results due to errors. A circuit partitioning approach can improve fidelity by separating and reducing the size of circuits processed at once in NISQ devices. In this paper, we propose Quixote (quantum independent execution architecture), that can execute quantum circuits independently as subcircuits to improve the fidelity of NISQ program results. We present methods for decomposing controlled gates into independent subcircuits and additional techniques for reducing circuit costs through identical gate transformation.
Enhyeok Jang, Seungwoo Choi 0001, Won Woo Ro
DAC1