Junyu Zhou 0005

dblp:29/4103-5 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0007-5564-8401ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 QTurbo: A Robust and Efficient Compiler for Analog Quantum Simulation
abstract
Analog quantum simulation leverages native hardware dynamics to emulate complex quantum systems with great efficiency by bypassing the quantum circuit abstraction. However, conventional compilation methods for analog simulators are typically labor-intensive, prone to errors, and computationally demanding. This paper introduces QTurbo, a powerful analog quantum simulation compiler designed to significantly enhance compilation efficiency and optimize hardware execution time. By generating precise and noise-resilient pulse schedules, our approach ensures greater accuracy and reliability, outperforming the existing state-of-the-art approach.
Junyu Zhou 0005, Yuhao Liu 0017, Shize Che, Anupam Mitra, Efekan Kökcü, Ermal Rrapaj, Costin Iancu, Gushu Li
ASPLOS (1)1
2026 AlphaSyndrome: Tackling the Syndrome Measurement Circuit Scheduling Problem for QEC Codes
abstract
Quantum error correction (QEC) is essential for scalable quantum computing, yet repeated syndrome-measurement cycles dominate its spacetime and hardware cost. Although stabilizers commute and admit many valid execution orders, different schedules induce distinct error-propagation paths under realistic noise, leading to large variations in logical error rate. Outside of surface codes, effective syndrome-measurement scheduling remains largely unexplored. We present AlphaSyndrome, an automated synthesis framework for scheduling syndrome-measurement circuits in general commuting-stabilizer codes under minimal assumptions: mutually commuting stabilizers and a heuristic decoder. AlphaSyndrome formulates scheduling as an optimization problem that shapes error propagation to (i) avoid patterns close to logical operators and (ii) remain within the decoder's correctable region. The framework uses Monte Carlo Tree Search (MCTS) to explore ordering and parallelism, guided by code structure and decoder feedback. Across diverse code families, sizes, and decoders, AlphaSyndrome reduces logical error rates by 80.6% on average (up to 96.2%) relative to depth-optimal baselines, matches Google's hand-crafted surface-code schedules, and outperforms IBM's schedule for the Bivariate Bicycle code.
Yuhao Liu 0017, Shuohao Ping, Junyu Zhou 0005, Ethan Decker, Justin Kalloor, Mathias Weiden, Kean Chen, Yunong Shi, Ali Javadi-Abhari, Costin Iancu, Gushu Li
ASPLOS (2)3
2026 Kernpiler: Compiler Optimization for Quantum Hamiltonian Simulation with Partial Trotterization
abstract
Description This artifact contains the core implementation of Kernpiler, a compiler framework for optimizing quantum circuits, and supports full reproducibility of all experimental results presented in the associated paper. The artifact includes all code, data pipelines, and scripts required to regenerate Figures 5–11. System Used for Data Collection NVIDIA A100 GPU with 80GB memory AMD EPYC 9654P 96-core processor x86_64 Linux system Python 3.13.5 Experiments may be computationally intensive, but they are fully parallelizable across multiple devices. Installation Clone or download the repository and navigate to the project directory. Create a virtual environment: python3 -m venv validate source validate/bin/activate Install dependencies: python -m pip install -r requirements.txt Install torch-scatter: python -m pip install --no-cache-dir torch-scatter -f https://data.pyg.org/whl/torch-2.10.0+cu128.html Experiment Workflow All experiment scripts are located in: src/compiler/optimization_passes/experiments Each figure can be reproduced by running its corresponding data collection and graphing scripts: Figure 5exp_gatecount_datacollection.py→ graph using:graph_data_scripts/graph_absolute.py Figure 6exp_partition_scaling_datacollection_o1exp_partition_scaling_datacollection→ graph using:graph_data_scripts/graph_o1_o2_side_by_side.py Figure 7exp_runtime_per_pass.py→ graph using:graph_data_scripts/graph_runtime_per_pass.py Figure 8exp_partition_scaling_datacollectiono1_phoenixFT→ graph using:graph_data_scripts/graph_firstorder_scalingFT.py Figure 9exp_partitionalgvsrandom.py→ graph using:graph_data_scripts/graph_partition_vs_random.py Figure 10exp_scaling_data_rewriteradius.py→ graph using:graph_data_scripts/graph_scalingdata.py Figure 11exp_error_scaling_systemsize.py→ output generated directly (no additional graph script required) Execution Notes All experiments are independent Parallel execution is supported Runtime varies depending on system size and hardware
Ethan Decker, Lucas Goetz, Evan McKinney, Erik Gustafson, Junyu Zhou 0005, Alex K. Jones, Ang Li 0006, Alexander Schuckert, Samuel A. Stein, Eleanor Crane, Gushu Li
ISCA5
2025 MarQSim: Reconciling Determinism and Randomness in Compiler Optimization for Quantum Simulation
abstract
Quantum Hamiltonian simulation, fundamental in quantum algorithm design, extends far beyond its foundational roots, powering diverse quantum computing applications. However, optimizing the compilation of quantum Hamiltonian simulation poses significant challenges. Existing approaches fall short in reconciling deterministic and randomized compilation, lack appropriate intermediate representations, and struggle to guarantee correctness. Addressing these challenges, we present MarQSim, a novel compilation framework. MarQSim leverages a Markov chain-based approach, encapsulated in the Hamiltonian Term Transition Graph, adeptly reconciling deterministic and randomized compilation benefits. Furthermore, we formulate a Minimum-Cost Flow model that can tune transition matrices to enforce correctness while accommodating various optimization objectives. Experimental results demonstrate MarQSim’s superiority in generating more efficient quantum circuits for simulating various quantum Hamiltonians while maintaining precision.
Xiuqi Cao, Junyu Zhou 0005, Yuhao Liu 0017, Yunong Shi, Gushu Li
Proc. ACM Program. Lang.2
2024 Fermihedral: On the Optimal Compilation for Fermion-to-Qubit Encoding
abstract
This paper introduces Fermihedral, a compiler framework focusing on discovering the optimal Fermion-to-qubit encoding for targeted Fermionic Hamiltonians. Fermion-to-qubit encoding is a crucial step in harnessing quantum computing for efficient simulation of Fermionic quantum systems. Utilizing Pauli algebra, Fermihedral redefines complex constraints and objectives of Fermion-to-qubit encoding into a Boolean Satisfiability problem which can then be solved with high-performance solvers. To accommodate larger-scale scenarios, this paper proposed two new strategies that yield approximate optimal solutions mitigating the overhead from the exponentially large number of clauses. Evaluation across diverse Fermionic systems highlights the superiority of Fermihedral, showcasing substantial reductions in implementation costs, gate counts, and circuit depth in the compiled circuits. Real-system experiments on IonQ's device affirm its effectiveness, notably enhancing simulation accuracy.
Yuhao Liu 0017, Shize Che, Junyu Zhou 0005, Yunong Shi, Gushu Li
ASPLOS (3)3
2024 Bosehedral: Compiler Optimization for Bosonic Quantum Computing
abstract
Bosonic quantum computing, based on the infinite-dimensional qumodes, has shown promise for various practical applications that are classically hard. However, the lack of compiler optimizations has hindered its full potential. This paper introduces Bosehedral, an efficient compiler optimization framework for (Gaussian) Boson sampling on Bosonic quantum hardware. Bosehedral overcomes the challenge of handling infinite-dimensional qumode gate matrices by performing all its program analysis and optimizations at a higher algorithmic level, using a compact unitary matrix representation. It optimizes qumode gate decomposition and logical-to-physical qumode mapping, and introduces a tunable probabilistic gate dropout method. Overall, Bosehedral significantly improves the performance by accurately approximating the original program with much fewer gates. Our evaluation shows that Bosehedral can largely reduce the program size but still maintain a high approximation fidelity, which can translate to significant end-to-end application performance improvement.
Junyu Zhou 0005, Yuhao Liu 0017, Yunong Shi, Ali Javadi-Abhari, Gushu Li
ISCA1