Yilun Zhao 0002

dblp:271/8391-2 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-6812-5120ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 SuperEncoder: Towards Efficient Neural Approximate Quantum State Preparation
abstract
Numerous quantum algorithms assume that classical data has already been converted into quantum states, a process known as Quantum State Preparation (QSP). However, achieving precise QSP requires a circuit depth that scales exponentially with the number of qubits, posing a significant challenge to realizing quantum advantage. Recent research explores Parameterized Quantum Circuits (PQCs) as an approximate alternative, offering improved scalability with reduced circuit depth. However, the iterative, state-by-state optimization required by this approach creates substantial runtime overhead, which severely limits its practicality.To improve the efficiency of approximate QSP, we introduce a novel two-stage framework that can potentially generate QSP circuits for arbitrary quantum states. In theoffline training stage, our model learns a direct mapping from target states to circuit parameters, thereby bypassing the need foronline, state-by-state optimizationduring theinference stage. Extensive evaluations show that our approach significantly reduces runtime overhead by up to 132×, making a steady step towards efficient neural approximate QSP.
Yilun Zhao 0002, Bingmeng Wang, Wenle Jiang, Xiwei Pan 0001, Bing Li 0017, Yinhe Han 0001, Ying Wang 0001
IEEE Trans. Computers1
2025 CLASS: A Controller-Centric Layout Synthesizer for Dynamic Quantum Circuits
abstract
Layout Synthesis for Quantum Computing (LSQC) is a critical component of quantum design tools. Traditional LSQC studies primarily focus on optimizing for reduced circuit depth by adopting a device-centric design methodology. However, these approaches overlook the impact of classical processing and communication time, thereby being insufficient for Dynamic Quantum Circuits (DQC).To address this, we introduce CLASS, a controller-centric layout synthesizer designed to reduce inter-controller communication latency in a distributed control system. It consists of a two-stage framework featuring a hypergraph-based modeling and a heuristic-based graph partitioning algorithm. Evaluations demonstrate that CLASS effectively reduces communication latency by up to 100% with only a 2.10% average increase in the number of additional operations.
Yilun Zhao 0002, Bing Li 0017, He Li 0008, Mengdi Wang 0004, Yinhe Han 0001, Ying Wang 0001
ICCAD2
2025 Distributed-HISQ: A Distributed Quantum Control Architecture
abstract
The design of a scalable Quantum Control Architecture (QCA) faces two primary challenges.First, the continuous growth in qubit counts has rendered distributed QCA inevitable, yet the nondeterministic latencies inherent in feedback loops demand cycleaccurate synchronization across multiple controllers.Existing synchronization strategies -whether lock-step or demand-drivenintroduce significant performance penalties.Second, existing quantum instruction set architectures are polarized, being either too abstract or too granular.This lack of a unifying design necessitates recurrent hardware customization for each new control requirement, which limits the system's reconfigurability and impedes the path toward a scalable and unified digital microarchitecture.Addressing these challenges, we propose Distributed-HISQ, featuring: (i) HISQ, A universal instruction set that redefines quantum control with a hardware-agnostic design.By decoupling from quantum operation semantics, HISQ provides a unified language for control sequences, enabling a single microarchitecture to support various control methods and enhancing system reconfigurability.(ii) BISP, a booking-based synchronization protocol that can potentially achieve zero-cycle synchronization overhead.The feasibility and adaptability of Distributed-HISQ are validated through its implementation on a commercial quantum control system targeting superconducting qubits.We performed a comprehensive evaluation using a customized quantum software stack.Our results show that BISP effectively synchronizes multiple control boards, leading to a 22.8% reduction in average program execution time and a ∼ 5× reduction in infidelity when compared to an existing lock-step synchronization scheme.
Yilun Zhao 0002, Kangding Zhao, Dingdong Liu, Tingyu Luo, Yuzhen Zheng, Shun Hu, Yinhe Han 0001, Ying Wang 0001, Mingtang Deng, Junjie Wu 0003, Xiang Fu 0003
MICRO1
2023 Full State Quantum Circuit Simulation Beyond Memory Limit
abstract
Quantum circuit simulation (QCS) is essential in the noisy intermediate scale quantum (NISQ) era when real quantum computers are scarce. However, fully tracking the states of a quantum system in QCS is highly challenging due to the exponential memory growth that significantly limits the computational reach of classical systems for QCS. Though it is straightforward to leverage secondary storage to extend the scale of QCS, excessive data movement between memory and storage dominates the simulation time, making this solution unrealistic. To tackle this challenge, we identify an intrinsic property of QCS and implement an open-source framework to effectively reduce data movement by >116x. We evaluate the framework on various benchmarks and demonstrate 4x memory reduction with only <20% overhead. On a memory constrained system, we show that it extends the scale of QCS to 32 qubits (64 GB memory requirement) while existing simulators are bounded to 28 qubits (4 GB memory requirement). Our implementation can be accessed via https://github.com/Zhaoyilunnn/qdao.
Yilun Zhao 0002, He Li 0008, Ying Wang 0001, Bingmeng Wang, Bing Li 0017, Yinhe Han 0001
ICCAD1
2022 Q-GPU: A Recipe of Optimizations for Quantum Circuit Simulation Using GPUs
abstract
In recent years, quantum computing has undergone significant developments and has established its supremacy in many application domains. While quantum hardware is accessible to the public through the cloud environment, a robust and efficient quantum circuit simulator is necessary to investigate the constraints and foster quantum computing development, such as quantum algorithm development and quantum device architecture exploration. In this paper, we observe that most of the publicly available quantum circuit simulators (e.g., QISKit from IBM, QDK from Microsoft, and Qsim-Cirq from Google) suffer from slow simulation and poor scalability when the number of qubits increases. To this end, we systematically investigate the deficiencies in quantum circuit simulation (QCS) and propose Q-GPU, a framework that leverages GPUs with comprehensive optimizations to allow efficient and scalable QCS. Specifically, Q-GPU features i) proactive state amplitude transfer, ii) zero state amplitude pruning, iii) delayed qubit involvement, and iv) lossless nonzero state amplitude compression. Experimental results across nine representative quantum circuits indicate that Q-GPU significantly reduces the execution time of the state-of-the-art GPU-based QCS by 71.89% (3.55× speedup). Q-GPU also outperforms the state-of-the-art OpenMP CPU implementation, the Google Qsim-Cirq simulator, and the Microsoft QDK simulator by 1.49×, 2.02×, and 10.82×, respectively.
Yilun Zhao 0002, Yanan Guo 0002, Amanda Dumi, Devin M. Mulvey, Shiv Upadhyay, Youtao Zhang, Kenneth D. Jordan, Jun Yang 0002, Xulong Tang
HPCA1