Jinglei Cheng

dblp:232/0135 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-9535-6672ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 EDDQC: Enhanced Dynamical Distributing Quantum Compilation
abstract
This article presents enhanced dynamical distributing quantum compilation (EDDQC), an optimized method for distributed quantum computing (DQC) using linear nearest neighbor (LNN) architecture integrated with quantum switches. By leveraging the symmetry of LNN topology and designating dangling qubits as communication links, our approach optimizes compilation for high local connectivity, sparse full connectivity algorithms (HLC-SFC) like quantum approximate optimization algorithm (QAOA) and quantum Fourier transform (QFT). Experimental results demonstrate significant performance improvements over traditional methods, including reductions in cross-group swaps by up to 67.2%, gate count by 43.8%, and total execution cycles by up to 40%. We also utilize the area law of entanglement entropy to limit entanglement growth in our 1-D system. Our comprehensive approach combines LNN chains’ efficiency with reconfigurable network flexibility, enhancing scalability and robustness for large-scale quantum computations.
Haochen Luo, Lingjun Xiong, Eilis Casey, Jinglei Cheng, Samuel Yen-Chi Chen, Zhiding Liang
IEEE Trans. Very Large Scale Integr. Syst.7
2025 Scalable Community Detection Using Quantum Hamiltonian Descent and QUBO Formulation
abstract
We present a quantum-inspired algorithm that utilizes Quantum Hamiltonian Descent (QHD) for efficient community detection. Our approach reformulates the community detection task as a Quadratic Unconstrained Binary Optimization (QUBO) problem, and QHD is deployed to identify optimal community structures. We implement a multi-level algorithm that iteratively refines community assignments by alternating between QUBO problem setup and QHD-based optimization. Benchmarking shows our method achieves up to 5.49% better modularity scores while requiring less computational time compared to classical optimization approaches. This work demonstrates the potential of hybrid quantum-inspired solutions for advancing community detection in largescale graph data.
Jinglei Cheng, Ruilin Zhou, Yuhang Gan, Chen Qian 0001, Junyu Liu
DAC1
2025 EPOC: An Efficient Pulse Generation Framework with Advanced Synthesis for Quantum Circuits
abstract
In this work, we aim to address the computational overhead challenge in quantum optimal control while reducing circuit latency. We propose a novel approach combining ZX-Calculus, circuit partitioning, and circuit synthesis for pulse generation. By implementing finer granularity in pulse generation and exploring equivalent circuit representations, we achieve increased parallelism and decreased latency. Our method demonstrates a 31.74% reduction in latency compared to previous work and a 76.80% reduction compared to gate-based pulse generation methods, while minimizing computational overhead.
Jinglei Cheng, Zhixin Song, Zhiding Liang
DAC1
2025 Hardware-aware Calibration Protocol for Quantum Computers
abstract
Calibration of a quantum computer is the process of optimizing its control parameters to ensure the accurate implementation of quantum gates.It remains a critical challenge in scaling quantum computers.Existing calibration methods take a generalized approach that focuses on the trade-off between calibration time and fidelity.However, these methods lack the awareness of hardware differences among physical qubits and an elaborate design of parallel calibration.In this paper, we introduce a fine-grained calibration protocol that contains three calibration policies for hardware differences and a method to enable parallel calibration.We begin by profiling qubit pairs to evaluate their responses to different waveform candidates.Based on profiling results, we determine the best calibration policy for the quantum computer, which is the first part of the calibration protocol.The second part of our protocol is to use graph traverse to enable parallel calibration by identifying compatible calibration operations.We validate our protocol through intensive experiments on real quantum machines with up to 127 qubits.Our experimental results demonstrate a 1.84× reduction in terms of the medium of the two-qubit gate error rate, 1.26× reduction in pulse duration, an 8× to 25× reduction in total calibration overhead compared with sequential calibration, an average of 2.12× further reduction in total calibration overhead owing to profiling policy, double of the quantum volume, and a 2.0× to 2.3× reduction in error per layered gate.The proposed protocol emphasizes the importance of hardware-aware and parallel calibration and advances current quantum computers towards fault-tolerant quantum computing.
Jinglei Cheng, Boxi Li, Hanrui Wang 0002, Yufei Ding 0001, Zhiding Liang
ISCA2
2025 ECDQC: Efficient Compilation for Distributed Quantum Computing with Linear Layout
abstract
In this paper, we propose an efficient compilation method for distributed quantum computing (DQC) using the Linear Nearest Neighbor (LNN) architecture. By exploiting the LNN topology’s symmetry, we optimize quantum circuit compilation for High Local Connectivity, Sparse Full Connectivity (HLC-SFC) algorithms like Quantum Approximate Optimization Algorithm (QAOA) and Quantum Fourier Transform (QFT). We also utilize dangling qubits to minimize non-local interactions and reduce SWAP gates. Our approach significantly decreases compilation time, gate count, and circuit depth, improving scalability and robustness for large-scale quantum computations.
Haochen Luo, Lingjun Xiong, Eilis Casey, Jinglei Cheng, Samuel Yen-Chi Chen, Zhiding Liang
ISCAS7
2024 Invited: Graph Learning for Parameter Prediction of Quantum Approximate Optimization Algorithm
abstract
In recent years, quantum computing has emerged as a transformative force in the field of combinatorial optimization, offering novel approaches to tackling complex problems that have long challenged classical computational methods. Among these, the Quantum Approximate Optimization Algorithm (QAOA) stands out for its potential to efficiently solve the Max-Cut problem, a quintessential example of combinatorial optimization. However, practical application faces challenges due to current limitations on quantum computational resource. Our work optimizes QAOA initialization, using Graph Neural Networks (GNN) as a warm-start technique. This sacrifices affordable computational resource on classical computer to reduce quantum computational resource overhead, enhancing QAOA's effectiveness. Experiments with various GNN architectures demonstrate the adaptability and stability of our framework, highlighting the synergy between quantum algorithms and machine learning. Our findings show GNN's potential in improving QAOA performance, opening new avenues for hybrid quantum-classical approaches in quantum computing and contributing to practical applications.
Zhiding Liang, Gang Liu 0025, Zheyuan Liu 0010, Jinglei Cheng, Tianyi Hao 0003, Zhixin Song, Ji Liu 0007, Fanny Ye, Yiyu Shi 0001
DAC4
2024 Combining Parameterized Pulses and Contextual Subspace for More Practical VQE
abstract
In this paper, we explore the integration of parameterized quantum pulses with the contextual subspace method. The advent of parameterized quantum pulses marks a transition from traditional quantum gates to a more flexible and efficient approach to quantum computing. Working with pulses allows us to potentially access areas of the Hilbert space that are inaccessible with a CNOT-based circuit decomposition. Compared to solving the complete Hamiltonian via the traditional Variational Quantum Eigensolver (VQE), the computation of the contextual correction generally requires fewer qubits and measurements, thus improving computational efficiency. Plus a Pauli grouping strategy, our framework, SpacePulse, can minimize the quantum resource cost for the VQE and enhance the potential for processing larger molecular structures.
Zhiding Liang, Zhixin Song, Jinglei Cheng, Tianyi Hao 0003, Yiyu Shi 0001, Tongyang Li
DAC3
2024 Compiler Optimizations for QAOA
abstract
The Quantum Approximate Optimization Algorithm (QAOA) is one of the most promising candidates for achieving quantum advantage over classical computers. However, existing compilers lack specialized methods for optimizing QAOA circuits. There are circuit patterns inside the QAOA circuits, and current quantum hardware has specific qubit connectivity topologies. Therefore, we propose Coqa to optimize QAOA circuit compilation tailored to different types of quantum hardware. Our method integrates a linear nearest-neighbor (LNN) topology and efficiently map the patterns of QAOA circuits to the LNN topology by heuristically checking the interaction based on the weight of problem Hamiltonian. This approach allows us to reduce the number of SWAP gates during compilation, which directly impacts the circuit depth and overall fidelity of the quantum computation. By leveraging the inherent patterns in QAOA circuits, our approach achieves more efficient compilation compared to general-purpose compilers. With our proposed method, we are able to achieve an average of 30% reduction in gate count and a 39x acceleration in compilation time across our benchmarks.
Jinglei Cheng, Yuwei Jin, Boxi Li, Siyuan Niu, Zhiding Liang
ICCAD3
2024 NAPA: Intermediate-Level Variational Native-Pulse Ansatz for Variational Quantum Algorithms
abstract
Variational quantum algorithms (VQAs) have demonstrated great potentials in the Noisy Intermediate Scale Quantum (NISQ) era. In the workflow of VQA, the parameters of ansatz are iteratively updated to approximate the desired quantum states. We have seen various efforts to draft better ansatz with less gates. Some works consider the physical meaning of the underlying circuits, while others adopt the ideas of neural architecture search (NAS) for ansatz generator. However, these designs do not exploit the full advantages of VQAs. Because most techniques target gate ansatz, and the parameters are usually rotation angles of the gates. In quantum computers, the gate ansatz will eventually be transformed into control signals such as microwave pulses on superconducting qubits. These control pulses need elaborate calibrations to minimize the errors such as over-rotation and under-rotation. In the case of VQAs, this procedure will introduce redundancy, but the variational properties of VQAs can naturally handle problems of over-rotation and under-rotation by updating the amplitude and frequency parameters. Therefore, we propose NAPA, a native-pulse ansatz generator framework for VQAs. We generate native-pulse ansatz with trainable parameters for amplitudes and frequencies. In our proposed NAPA, we are tuning parametric pulses, which are natively supported on NISQ computers. Given the limited availability of gradient-based optimizers for pulse-level quantum programs, we choose to deploy non-gradient optimizers in our framework. To constrain the number of parameters sent to the optimizer, we adopt a progressive way to generate our nativepulse ansatz. Experiments are conducted on both simulators and quantum devices for Variational Quantum Eigensolver (VQE) tasks to envaluate our methods. When adopted on NISQ machines, NAPA obtained improved the performance with decreased latency by an average of 86%. NAPA is able to achieve 96.482% and 99.336% accuracy for VQE tasks on H2 and HeH+ respectively. An average accuracy of 97.27% is achieved for medium-size quantum chemistry tasks on CO2, H2O, and NaH. NAPA also demonstrates advantages on quantum optimization tasks even with considerable noises in NISQ machines.
Zhiding Liang, Jinglei Cheng, Hanrui Wang 0002, Zhixin Song, Yongshan Ding 0001, Fred Chong, Song Han 0003, Xuehai Qian, Yiyu Shi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Hybrid Gate-Pulse Model for Variational Quantum Algorithms
abstract
Current quantum programs are mostly synthesized and compiled on the gate-level, where quantum circuits are composed of quantum gates. The gate-level workflow, however, introduces significant redundancy when quantum gates are eventually transformed into control signals and applied on quantum devices. For superconducting quantum computers, the control signals are microwave pulses. Therefore, pulse-level optimization has gained more attention from researchers due to their advantages in terms of circuit duration. Recent works, however, are limited by their poor scalability brought by the large parameter space of control signals. In addition, the lack of gate-level "knowledge" also affects the performance of pure pulse-level frameworks. We present a hybrid gate-pulse model that can mitigate these problems. We propose to use gate-level compilation and optimization for "fixed" part of the quantum circuits and to use pulse-level methods for problem-agnostic parts. Experimental results demonstrate the efficiency of the proposed framework in discrete optimization tasks. We achieve a performance boost at most 8% with 60% shorter pulse duration in the problem-agnostic layer.
Zhiding Liang, Zhixin Song, Jinglei Cheng, Zichang He, Ji Liu 0007, Hanrui Wang 0002, Ruiyang Qin, Song Han 0003, Xuehai Qian, Yiyu Shi 0001
DAC3
2020 AccQOC: Accelerating Quantum Optimal Control Based Pulse Generation
abstract
In the last decades, we have witnessed the rapid growth of Quantum Computing. In the current Noisy Intermediate-Scale Quantum (NISQ) era, the capability of a quantum machine is limited by the decoherence time, gate fidelity and the number of Qubits. Current quantum computing applications are far from the real “quantum supremacy” due to the fragile physical Qubits, which can only be entangled for a few microseconds. Recent works use quantum optimal control to reduce the latency of quantum circuits, thereby effectively increasing quantum volume. However, the key challenge of this technique is the large overhead due to long compilation time. In this paper, we propose AccQOC, a comprehensive static/dynamic hybrid workflow to transform gate groups (equivalent to matrices) to pulses using QOC (Quantum Optimal Control) with a reasonable compilation time budget. AccQOC is composed of static pre-compilation and accelerated dynamic compilation. After the quantum program is mapped to the quantum circuit with our heuristic mapping algorithm considering crosstalk, we leverage static pre-compilation to generate pulses for the frequently used groups to eliminate the dynamic compilation time for them. The pulse is generated using QOC with binary search to determine the latency. For a new program, we use the same policy to generate groups, thus avoid incurring overhead for the “covered” groups. The dynamic compilation deals with “un-covered” groups with accelerated pulse generation. The key insight is that the pulse of a group can be generated faster based on the generated pulse of a similar group. We propose to reduce the compilation time by generating an ordered sequence of groups in which the sum of similarity among consecutive groups in the sequence is minimized. We can find the sequence by constructing a similarity graph - a complete graph in which each vertex is a gate group and the weight of an edge is the similarity between the two groups it connects, then construct a Minimum Spanning Tree (MST) for SG. With the methodology of AccQOC, we reached a balanced point of compilation time and overall latency. The results show that accelerated compilation based on MST achieves 9.88× compilation speedup compared to the standard compilation of each group while maintaining an average 2.43× latency reduction compared with gate-based compilation.
Jinglei Cheng, Haoqing Deng, Xuehai Qian
ISCA1
2018 CSE: Parallel Finite State Machines with Convergence Set Enumeration
abstract
Finite State Machine (FSM) is known to be “embarrassingly sequential” because the next state depends on the current state and input symbol. Enumerative FSM breaks the data dependencies by cutting the input symbols into segments and processing all segments in parallel. With unknown starting state (except the first segment), each segment needs to calculate the state transitions, i.e., state state, for all states, each one is called an enumeration path. The current software and hardware implementations suffer from two drawbacks: 1) large amount of state state computation overhead for the enumeration paths; and 2) the optimizations are restricted by the need to correctly performing state state and only achieve limited improvements. This paper proposes CSE, a Convergence Set based Enumeration based parallel FSM. Unlike prior approaches, CSE is based on a novel computation primitive set(N) set(M), which maps N states to M states without giving the specific state state mappings (which state is mapped to which). The set(N) set(M) has two key properties: 1) if M is equal to 1, i.e., all N states are mapped to the same state, the state state for all the N states are computed; 2) using one-hot encoding, the hardware implementation cost of state state is the same as set(N) set(M). The convergence property ensures that M is always less than N. The key idea of CSE is to partition the original all S states into n state sets CS1,CS2,...,CSn, i.e., convergence sets. Using set(N) set(M) to process each CS, if the states converge to a single state, then we have successfully computed the enumeration path for each state in CS; otherwise, we may need to re-execute the stage when the outcome of the previous stage falls in CS. CSE is realized by two techniques: convergence set prediction, which generates the convergence sets with random input based profiling that maximizes the probability of each CS z converging to one state; global re-execution algorithm, which ensures the correctness by re-executing the non-converging stages with known input state. Essentially, CSE reformulates the enumeration paths as setbased rather than singleton-based. We evaluate CSE with 13 benchmarks. It achieved on average 2.0x/2.4x and maximum 8.6x/2.7x speedup compared to Lookback Enumeration (LBE) and Parallel Automata Processor (PAP), respectively.
Youwei Zhuo, Jinglei Cheng, Qinyi Luo, Jidong Zhai, Yanzhi Wang 0001, Zhongzhi Luan, Xuehai Qian
MICRO2