Kaiwen Zhou 0003

dblp:215/4936-3 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0005-1017-8878ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 DyREM: Dynamically Mitigating Quantum Readout Error with Embedded Accelerator
abstract
Quantum readout error is the most significant source of error, substantially reducing the measurement fidelity. Tensor-product-based readout error mitigation has been proposed to address this issue by approximating the mitigation matrix. However, this method inevitably encounters the dynamic generation of the mitigation matrix, leading to long latency. In this paper, we propose DyREM, a software-hardware codesign approach that mitigates readout errors with an embedded accelerator. The main innovation lies in leveraging the inherent sparsity in the nonzero probability distribution of quantum states and calculating the tensor product on an embedded accelerator. Specifically, using the output sparsity, our dataflow dynamically downsamples the original mitigation matrix, which dramatically reduces the memory requirement. Then, we design DyREM architecture that can flexibly gate the redundant computation of nonzero quantum states. Experiments demonstrate that DyREM achieves an average speedup of $9.6 \times \sim 2000 \times$ and fidelity improvements of $1.03 \times \sim 1.15 \times$ compared to state-of-the-art readout error mitigation methods.
Kaiwen Zhou 0003, Liqiang Lu, Debin Xiang, Chenning Tao, Xinkui Zhao, Size Zheng 0001, Jianwei Yin
DAC1
2025 Qtenon: Towards Low-Latency Architecture Integration for Accelerating Hybrid Quantum-Classical Computing
abstract
Hybrid quantum-classical algorithms have shown great promise in leveraging the computational potential of quantum systems.However, the efficiency of these algorithms is severely constrained by the limitations of current quantum hardware architectures.These architectures, which typically feature a decoupled design, lack both hardware support for low-latency communication and software support for fine-grained optimization.In this paper, we propose Qtenon, a tightly coupled system for efficient hybrid quantum-classical algorithm acceleration.Qtenon is composed of both hardware part and software part.To enable efficient communication and computation, the hardware part provides a unified memory hierarchy, an efficient quantum controller, as well as a multi-stage processing pipeline.The unified memory hierarchy functions as a communication buffer between host and quantum accelerators, with dedicated data paths and interfaces provided by the quantum controller.The multi-stage pipeline leverages hardware pipelines to fully exploit parallelism.To program hybrid quantum-classical algorithms on the hardware, our software part provides a set of instructions for data communication and computation.The instructions also enable fine-grained synchronization and efficient scheduling for quantum-host interaction.We design Qtenon as a RISC-V extended chip and implement it using Chisel.In evaluation, we achieve up to 14.9× end-to-end speedup compared to state-of-the-art work for hybrid quantum-classical algorithms.
Chenning Tao, Liqiang Lu, Size Zheng 0001, Li-Wen Chang, Minghua Shen, Fangxin Liu, Kaiwen Zhou 0003, Jianwei Yin
ISCA8
2025 ARTERY: Fast Quantum Feedback using Branch Prediction
abstract
Quantum feedback makes the execution of dynamic quantum circuits possible and is widely used in quantum algorithms.However, due to the inherent computation and transmission cost, the latency of the quantum feedback becomes a considerable burden on the current quantum algorithm.The dynamic property of the feedback also makes the gates blocked until the feedback is finished.In this paper, we propose ARTERY, which uses branch prediction to support instruction pre-execution and speed up the feedback.ARTERY integrates historical statistics of branches and a real-time readout pulse analysis to predict the branch.With this idea, we build up a reconciled branch predictor that concatenates the historical statistics of branches and a real-time branch circuit speculation obtained from the readout-pulse trajectory predictor.We further explore the implementation of peripheral hardware for feedback, including a scalable inter-FPGA connection via the backplane, a feedback trigger mechanism for dynamic instruction timing, and an adaptive pulse sampling technique to maximize the hardware bandwidth.ARTERY accelerates quantum feedback process by 2.07× compared to the state-of-the-art method, with over 90% prediction accuracy, achieving 1.24× fidelity improvement.
Wuwei Tian, Liqiang Lu, Siwei Tan, Yun Liang 0001, Tingting Li 0004, Kaiwen Zhou 0003, Xinghui Jia, Jianwei Yin
ISCA6
2025 Vegapunk: Accurate and Fast Decoding for Quantum LDPC Codes with Online Hierarchical Algorithm and Sparse Accelerator
abstract
Quantum Low-Density Parity-Check (qLDPC) codes are a promising class of quantum error-correcting codes that exhibit constantrate encoding and high error thresholds, thereby facilitating scalable fault-tolerant quantum computation.However, real-time decoding of qLDPC codes remains a significant challenge due to the high connectivity of their check matrices, which typically requires solving large-scale linear systems with sparse structures.In particular, off-the-shelf qLDPC decoders are often subject to a tradeoff between accuracy and latency, thus yielding no accurate and realtime decoding.This paper presents Vegapunk, a software-hardware co-design framework that enables real-time qLDPC decoding with high accuracy.To improve decoding accuracy, we design an offline decoupling strategy leveraging Satisfiability Modulo Theories (SMT) optimizations to mitigate quantum degeneracy.To enable fast decoding, we introduce an online hierarchical decoding algorithm employing a greedy strategy.Furthermore, we show that our SMT-optimized strategy suffices to produce decoupled matrices with maximized sparsity, thus admitting a dedicated accelerator to fully exploit the sparsity and parallelism to achieve real-time qLDPC decoding.Experimental results demonstrate that Vegapunk enables real-time decoding (< 1𝜇𝑠) for the Bivariate Bicycle (BB) code up to [[784,24,24]] while exhibiting logical error rates on par with the state-of-the-art decoder, i.e., BP+OSD.
Kaiwen Zhou 0003, Liqiang Lu, Debin Xiang, Chenning Tao, Anbang Wu, Jingwen Leng, Fangxin Liu, Mingshuai Chen, Jianwei Yin
MICRO1