Yuwei Jin

dblp:125/0848 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Leveraging Phase Polynomials for Quantum Circuit Optimization
abstract
Quantum circuits on resource-limited hardware require optimizing regions dominated by $\{\mathrm{CNOT}, R_z\}$, which account for a large fraction of operations and often dominate execution cost. This optimization can be challenging because phase-polynomial blocks are fragmented by basis-changing gates such as $H$, and optimizing phase parities alone may increase the cost of downstream basis transformations. Existing phase-polynomial approaches are limited to single-block or phase-only optimization, while subcircuit rewriting approaches are local and scale poorly beyond small rewrite windows. We introduce \emph{PhasePoly}, a compiler optimization pass that jointly optimizes phase-parity and output-parity networks and employs a cross-block intermediate representation to reuse parities across phase-polynomial block barriers. This approach is effective because its unified parity-matrix representation exposes long-range $\{\mathrm{CNOT}, R_z\}$ structure that local rewriting and single-block methods cannot capture. \emph{PhasePoly} reduces total gate count by up to 50.00\% (34.70\% on average) and CNOT count by up to 48.57\% (26.83\% on average), while scaling to large circuits and improving both fault-tolerant compilation and near-term hardware execution. \emph{PhasePoly} is available at https://github.com/ruadapt/PhasePoly.
Zihan Chen 0005, Henry Chen, Yuwei Jin, Enhyeok Jang, Mingkuan Xu, Vannessa Chan, Won Woo Ro, Eddy Z. Zhang
ISCA3
2024 Compiler Optimizations for QAOA
abstract
The Quantum Approximate Optimization Algorithm (QAOA) is one of the most promising candidates for achieving quantum advantage over classical computers. However, existing compilers lack specialized methods for optimizing QAOA circuits. There are circuit patterns inside the QAOA circuits, and current quantum hardware has specific qubit connectivity topologies. Therefore, we propose Coqa to optimize QAOA circuit compilation tailored to different types of quantum hardware. Our method integrates a linear nearest-neighbor (LNN) topology and efficiently map the patterns of QAOA circuits to the LNN topology by heuristically checking the interaction based on the weight of problem Hamiltonian. This approach allows us to reduce the number of SWAP gates during compilation, which directly impacts the circuit depth and overall fidelity of the quantum computation. By leveraging the inherent patterns in QAOA circuits, our approach achieves more efficient compilation compared to general-purpose compilers. With our proposed method, we are able to achieve an average of 30% reduction in gate count and a 39x acceleration in compilation time across our benchmarks.
Jinglei Cheng, Yuwei Jin, Boxi Li, Siyuan Niu, Zhiding Liang
ICCAD4
2024 Tetris: A Compilation Framework for VQA Applications in Quantum Computing
abstract
Quantum computing has shown promise in solving complex problems by leveraging the principles of superposition and entanglement. Variational quantum algorithms (VQA) are a class of algorithms suited for near-term quantum computers due to their modest requirements of qubits and depths of computation. This paper introduces Tetris – a compilation framework for VQA applications on near-term quantum devices. Tetris focuses on reducing two-qubit gates in the compilation process since a two-qubit gate has an order of magnitude more significant error and execution time than a single-qubit gate. Tetris exploits unique opportunities in the circuit synthesis stage often overlooked by the state-of-the-art VQA compilers for reducing the number of two-qubit gates. Tetris comes with a refined IR of Pauli string to express such a two-qubit gate optimization opportunity. Moreover, Tetris is equipped with a fast bridging approach that mitigates the hardware mapping cost. Overall, Tetris demonstrates a reduction of up to $41.3 \%$ in CNOT gate counts, $37.9 \%$ in circuit depth, and $\mathbf{4 2. 6 \%}$ in circuit duration for various molecules of different sizes and structures compared with the state-of-the-art approaches. Tetris is open-sourced at this link.
Yuwei Jin, Tianyi Hao 0003, Huiyang Zhou, Yipeng Huang 0001, Eddy Z. Zhang
ISCA1
2024 Optimizing Quantum Fourier Transformation (QFT) Kernels for Modern NISQ and FT Architectures
abstract
Rapid development in quantum computing leads to the appearance of several quantum applications. Quantum Fourier Transformation (QFT) sits at the heart of many of these applications. Existing work leverages SAT solver or heuristics to generate a hardware-compliant circuit for QFT by inserting SWAP gates to remap logical qubits to physical qubits. However, they might face problems such as long compilation time due to the huge search space for SAT solver or suboptimal outcome in terms of the number of cycles to finish all gate operations. In this paper, we propose a domain-specific hardware mapping approach for QFT. We unify our insight of relaxed ordering and unit exploration in QFT to search for a qubit mapping solution with the help of program synthesis tools. Our method is the first one that guarantees linear-depth QFT circuits for Google Sycamore, IBM heavy-hex, and the lattice surgery, with respect to the number of qubits. Compared with state-of-the-art approaches, our method can save up to 53% in SWAP gate and 92% in depth.
Yuwei Jin, Henry Chen, Chi Zhang 0041, Eddy Z. Zhang
SC1
2024 Estimating Broadband Snow Albedo by an Asymptotic Radiative Transfer Model-Constrained Deep Learning Models
abstract
Deep learning (DL) has been widely applied to enhance the accuracy and reconstruction of snowpack parameters due to its powerful nonlinear mapping ability. Here, we propose a broadband snow albedo (BA) estimation method that infuses the asymptotic radiative transfer (ART) model into the DL model and thus imposes constraints on the input variables of the DL model and compares it to the MOD10A1 BA product and the BA simulated using the ART model alone. The results show that coupling a physical mechanism-based snow radiative transfer model with a DL model can successfully improve the estimation accuracy of BA. Compared to the MOD10A1 BA product and the BA simulated using the ART model alone, the accuracy of the BA simulated by the convolutional neural networks (CNNs) model under the constraints of the ART model is improved by 12% and 3%, respectively, and the accuracy of the BA simulated by the long short-term memory (LSTM) model under the constraints of the ART model is improved by 14% and 5%, respectively. The coupling paradigm that integrates the snow radiative transfer model into the DL model effectively balances the advantages of both models. This coupling paradigm creates a complementary association between the snow radiative transfer model and the DL model. This research effectively enhances the simulation and predictive capabilities of snow albedo and is expected to promote the in-depth application of remote sensing big data in combination with physical modeling in the field of cryosphere parameter estimation.
Donghang Shao, Yuwei Jin, Wenzheng Ji, Hongxing Li 0005, Tianwen Feng, Hongyi Li 0003, Jian Wang 0032, Xiaohua Hao
IEEE Trans. Geosci. Remote. Sens.2
2023 CaQR: A Compiler-Assisted Approach for Qubit Reuse through Dynamic Circuit
abstract
Quantum measurement is important to quantum computing as it extracts out the outcome of the circuit at the end of the computation. Previously, all measurements have to be done at the end of the circuit. Otherwise, it will incur significant errors. But it is not the case now. Recently IBM starts supporting dynamic circuit through hardware (instead of software by simulator). With mid-circuit hardware measurement, we can improve circuit efficacy and fidelity from three aspects: (a) reduced qubit usage, (b) reduced swap insertion, and (c) improved fidelity. We demonstrate this using real-world applications Bernstein Verizani on real hardware and show that circuit resource usage can be improved by 60%, and circuit fidelity can be improved by 15%. We design a compiler-assisted tool that can find and exploit the tradeoff between qubit reuse, fidelity, gate count, and circuit duration. We also developed a method for identifying whether qubit reuse will be beneficial for a given application. We evaluated our method on a representative set of important applications. We can reduce resource usage by up to 80% and improve circuit fidelity by up to 20%.
Yuwei Jin, Yan-Hao Chen, Suhas Vittal, Kevin Krsulich, Lev S. Bishop, John Lapeyre, Ali Javadi-Abhari, Eddy Z. Zhang
ASPLOS (3)2
2023 Exploiting the Regular Structure of Modern Quantum Architectures for Compiling and Optimizing Programs with Permutable Operators
abstract
A critical feature in today's quantum circuit is that they have permutable two-qubit operators. The flexibility in ordering the permutable two-qubit gates leads to more compiler optimization opportunities. However, it also imposes significant challenges due to the additional degree of freedom. Our Contributions are two-fold. We first propose a general methodology that can find structured solutions for scalable quantum hardware. It breaks down the complex compilation problem into two sub-problems that can be solved at small scale. Second, we show how such a structured method can be adapted to practical cases that handle sparsity of the input problem graphs and the noise variability in real hardware. Our evaluation evaluates our method on IBM and Google architecture coupling graphs for up to 1,024 qubits and demonstrate better result in both depth and gate count - by up to 72% reduction in depth, and 66% reduction in gate count. Our real experiments on IBM Mumbai show that we can find better expected minimal energy than the state-of-the-art baseline.
Yuwei Jin, Yan-Hao Chen, Ari B. Hayes, Chi Zhang 0041, Eddy Z. Zhang
ASPLOS (4)1
2023 A Pulse Generation Framework with Augmented Program-aware Basis Gates and Criticality Analysis
abstract
Near-term intermediate-scale quantum (NISQ) devices are subject to considerable noise and short coherence time. Consequently, it is critical to minimize circuit execution latency and improve fidelity. Traditionally, each basis gate of a transpiled circuit is decoded into a fixed episode of the device control pulses. Recent studies investigate the merged pulse generation method for customized gates through quantum optimal control (QOC). In this work, we propose PAQOC, a novel QOC framework that can (i) exploit an augmented program-aware (APA) basis gate set for the tradeoff between compilation time and circuit performance, (ii) prune the search space based on a criticality-centric analytical model and experiment observations we learned from 150 benchmarks. Evaluations using seventeen applications show that PAQOC can achieve an average 54% reduction of the circuit latency, on average 43% reduction in compilation overhead, and a 1.27× improvement in fidelity. PAQOC is available on GitHub1.
Yan-Hao Chen, Yuwei Jin, Ari B. Hayes, Ang Li 0006, Yunong Shi, Eddy Z. Zhang
HPCA2
2021 Time-optimal Qubit mapping
abstract
Rapid progress in the physical implementation of quantum computers gave birth to multiple recent quantum machines implemented with superconducting technology. In these NISQ machines, each qubit is physically connected to a bounded number of neighbors. This limitation prevents most quantum programs from being directly executed on quantum devices. A compiler is required for converting a quantum program to a hardware-compliant circuit, in particular, making each two-qubit gate executable by mapping the two logical qubits to two physical qubits with a link between them. To solve this problem, existing studies focus on inserting SWAP gates to dynamically remap logical qubits to physical qubits. However, most of the schemes lack the consideration of time-optimality of generated quantum circuits, or are achieving time-optimality with certain constraints. In this work, we propose a theoretically time-optimal SWAP insertion scheme for the qubit mapping problem. Our model can also be extended to practical heuristic algorithms. We present exact analysis results by using our model for quantum programs with recurring execution patterns. We have for the first time discovered an optimal qubit mapping pattern for quantum fourier transformation (QFT) on 2D nearest neighbor architecture. We also present a scalable extension of our theoretical model that can be used to solve qubit mapping for large quantum circuits.
Chi Zhang 0041, Ari B. Hayes, Longfei Qiu, Yuwei Jin, Yan-Hao Chen, Eddy Z. Zhang
ASPLOS4
2021 BGPQ: A Heap-Based Priority Queue Design for GPUs
abstract
Programming today’s many-core processor is challenging. Due to the enormous amount of parallelism, synchronization is expensive. We need efficient data structures for providing automatic and scalable synchronization methods. In this paper, we focus on the priority queue data structure. We develop a heap-based priority queue implementation called BGPQ. BGPQ uses batched key nodes as the internal data representation, exploits both task parallelism and data parallelism, and is linearizable. We show that BGPQ achieves up to 88X speedup compared with four state-of-the-art CPU parallel priority queue implementations and up to 11.2X speedup over an existing GPU implementation. We also apply BGPQ to search problems, including 0-1 Knapsack and A* search. We achieve 45X-100X and 12X-46X speedup respectively over best known concurrent CPU priority queues.
Yan-Hao Chen, Yuwei Jin, Eddy Z. Zhang
ICPP3
2021 AutoBraid: A Framework for Enabling Efficient Surface Code Communication in Quantum Computing
abstract
Quantum computers can solve problems that are intractable using the most powerful classical computer. However, qubits are fickle and error prone. It is necessary to actively correct errors in the execution of a quantum circuit. Quantum error correction (QEC) codes are developed to enable fault-tolerant quantum computing. With QEC, one logical circuit is converted into an encoded circuit.
Yan-Hao Chen, Yuwei Jin, Chi Zhang 0041, Ari B. Hayes, Youtao Zhang, Eddy Z. Zhang
MICRO3
2019 Object-Oriented Automatic and Accurate Shadow Detection for Very High Spatial Resolution Satellite Images
abstract
Several existing shadow detection methods cannot keep the balance between accuracy and automaticity well. To overcome the weakness, we present a novel method to detect shadow in very high spatial resolution satellite images. First, a new shadow detection index is developed to obtain the shadow ratio map. The initial shadow mask map is then obtained by utilizing the Gaussian mixture mode and the Otsu's method automatically. Finally, the initial shadow mask map is refined by jointly using the object spectral characteristics and the spatial-correlation relationship between objects. The experimental results performed on different images show that the accuracy and automation of the proposed method are over several state-of-the-art methods.
Yuwei Jin, Wenbo Xu 0004, Donghang Shao, Xixu He, Xueru Zhang
IGARSS1
2019 Forward Simulation of Snow Albedo Based on Snicar Model
abstract
Snow albedo plays an important role in the global climate system due to its climatic feedback effects. Due to the limitations of remote sensing methods, the remotely sensed snow albedo products have significant data loss and error uncertainty. In view of the limitation of retrieval methods, this research utilizing SNICAR model to study the forward simulation of snow albedo. We verify and optimize the SNICAR forward simulation model at the plot scale, based on field measurements of the input variables of forward model such as snow grain size, snow density, water content and solar zenith angle, and combined with the snow grain size evolution model driven by meteorological elements. The results indicate that the MAE (mean absolute error), RMSE (root mean square error), R (Pearson correlation coefficient), and NSE (Nash-Sutcliffe efficiency coefficient) of the observed and simulated snow albedo by the optimized SNICAR model are 0.04, 0.05, 0.86 and 0.67, respectively. Our research validates and optimizes the snow albedo forward simulation model, which provides an effective simulation means acquiring the snow albedo data of continuous time series in alpine mountain regions.
Donghang Shao, Wenbo Xu 0004, Hongyi Li 0003, Jian Wang 0032, Xiaohua Hao, Yuwei Jin
IGARSS7
2019 Extracting Land Surface Water from FY/MERSI Image Based On Spectral Matching Of Discrete Particle Swarm Optimization and Linear Feature Enhancement
abstract
Land surface water is one of the most important components of surface cover and global water cycle. In this study, the standard water spectrum selected from FY/MERSI image was firstly used to calculate the water probability. Then, based on water probability, the image was roughly classified into four classes: homogeneous ground, junction of land cover, minor tributaries and other. To extract land surface water from small tributaries, the Duda's Road Operator (DRO) was introduced to enhance the linear features, while Discrete Particle Swarm Optimization (DPSO) was applied to extract land surface water from other three classes. The results show that the method could effectively extract land surface water, especially from small tributaries, and overall accuracy (OA) and Kappa coefficient are improved compared to DPSO algorithm based on spectral matching (SMDPSO).
Xueru Zhang, Wenbo Xu 0004, Jinsheng Ren, Xixu He, Yuwei Jin
IGARSS7