Rongchao Dong

dblp:278/6131 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0004-5559-7075ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 STMC: Small-Tile Multiple-Copy Compilation for Reliable Measurement-Based Quantum Computing
abstract
Measurement-based Quantum Computing (MBQC) achieves universal quantum computing by applying measurements on the photonic architectures. While it has many advantages, such as long qubit decoherence time and strong scalability, the success rate of MBQC execution is constrained by imperfect photon control, measurement, and fusion operations. Both fusion failure and photon loss necessitate the re-execution of the entire quantum circuit, leading to significant overhead in terms of additional execution time and increased consumption of resource state layers. Recent studies mainly focus on mitigating fusion failures and little attention has been paid to photon loss. In this paper, we propose STMC (Small-Tile Multiple-Copy) compilation framework to reduce the re-execution overhead caused by both the fusion failure and photon loss. Specifically, STMC first transforms a quantum circuit into a fusion graph and partitions the fusion graph into subgraphs. Then, STMC generates compact subgraph mappings that are appropriate for the size of a subportion in the resource state layer, referred to as a tile. Finally, STMC employs multiple copies of each subgraph when mapping to tiles, duplicates the execution of tiles in parallel, and finishes the whole circuit execution in order. The experimental results demonstrate that STMC achieves an average execution time speedup of 65.68× for successfully executing the circuit under a 75% fusion success rate, compared to prior work. Additionally, STMC reduces the number of resource state layers by three orders of magnitude and decreases the number of resource states by an average of 36.40×.
Rongchao Dong, Zewei Mo, Yingheng Li, Aditya Pawar, Jun Yang 0002, Youtao Zhang, Xulong Tang
ICCAD1
2025 Reinforcement Learning-Guided Graph State Generation in Photonic Quantum Computers
abstract
The photonic quantum computer (PQC) is an emerging and promising quantum computing paradigm that has gained momentum in recent years.In PQC, computations are executed by performing measurements on photons in graph states (i.e., a collection of entangled photons).The graph state generation process is fulfilled by applying a sequence of quantum gates to quantum emitters, referred to as the "generation sequence".In a generation sequence, i) the time required to complete the generation sequence, ii) the number of quantum emitters used, and iii) the number of CZ gates performed between emitters greatly affect the fidelity of the generated graph state.In this paper, we propose RLGS (Reinforcement Learningguided Graph State generation), a novel compilation framework to identify optimal generation sequences that optimize the three fidelity metrics.Experimental results show that RLGS achieves an average reduction in generation time of 31.1%,49.6%, and 57.5% for small, medium, and large graph states compared to the baseline.Additionally, the reductions in the number of quantum emitters are 13.9%, 16.7%, and 17.5%, whereas the reductions in the number of CZ gates are 37.7%, 53.4%, and 57.8%, respectively.
Yingheng Li, Yue Dai 0005, Aditya Pawar, Rongchao Dong, Jun Yang 0002, Youtao Zhang, Xulong Tang
ISCA4
2024 A System-Level Dynamic Binary Translator Using Automatically-Learned Translation Rules
abstract
System-level emulators have been used extensively for the design, debugging and evaluation of the system software. They work by providing a system-level virtual machine that can support a guest operating system (OS) running on a platform with the same or different native OS using the same or different instruction-set architecture. For such a system-level emulation, dynamic binary translation (DBT) is one of the core technologies. A recently proposed learning-based approach using automatically-learned translation rules has shown to improve DBT performance significantly with much higher quality translated code. However, it has only been used on user-level emulation, not system-level emulation. In applying this approach directly on QEMU for system-level emulation, we find it actually causes an unexpected performance degradation of 5% on average. By analyzing its main culprits in more detail, we find that the learning-based approach will by default use host registers to maintain the guest CPU states that include condition-code registers (or FLAG registers). In cases where QEMU needs to be involved (in which QEMU also needs to use the host registers), maintaining system states in the host registers for the guest, the host and QEMU during and between the context switches can cause undue overheads, if not handled carefully. Such cases include emulating system-level instructions, address translation and interrupts, which require the use of QEMU's helper functions. To achieve the intended performance improvement through better-quality code generated by the learning-based approach, we propose several optimization techniques that include reducing the overhead incurred in each context switch, the number of needed context switches, and better code scheduling to eliminate context switches. Our experimental results show that such optimizations can achieve an average of 1.36X speedup over QEMU 6.1 using SPEC CINT2006 and 1.15X on real-world applications in the system emulation mode.
Jinhu Jiang, Chaoyi Liang, Rongchao Dong, Zhaohui Yang 0001, Zhongjun Zhou, Wenwen Wang 0001, Pen-Chung Yew
CGO3
2024 MixQ: Taming Dynamic Outliers in Mixed-Precision Quantization by Online Prediction
abstract
Mixed-precision quantization has shown to be a promising method for enhancing the efficiency of LLMs. This technique boosts computational efficiency by processing most values with low-precision, high-throughput compute units and maintains accuracy by processing outliers in high-precision. However, due to the dynamic, irregular, and sparse nature of outliers, this approach is far from using hardware efficiently. In this work, we propose MixQ, an efficient mixed-precision quantization system. Through our in-depth analysis of outlier distribution, we introduce a locality-based outlier prediction algorithm that can predict all outliers of 95.8% of tokens. Based on this accurate prediction, we propose a quantization ahead of detection (QAD) technique that can verify the correctness of prediction. A new data structure is proposed for efficient outlier processing. Evaluation shows that MixQ achieves $1.52 \times$ and $1.78 \times$ speedup over FP16 and Bitsandbytes on 8-bit quantization; plus $1.48 \times 1.93 \times$ and $6 \times$ speedup over QUIK, FP16, and AWQ on 4-bit quantization.11Our code is available on:https://github.com/Qcompiler/MIXQ
Yidong Chen 0003, Chen Zhang 0001, Rongchao Dong, Zhonghua Lu, Jidong Zhai
SC3
2024 MAGPY: Compiling Eager Mode DNN Programs by Monitoring Execution States
Chen Zhang 0001, Rongchao Dong, Haojie Wang 0004, Runxin Zhong, Jike Chen, Jidong Zhai
USENIX ATC2
2020 More with Less - Deriving More Translation Rules with Less Training Data for DBTs Using Parameterization
abstract
Dynamic binary translation (DBT) is widely used in system virtualization and many other important applications. To achieve a higher translation quality, a learning-based approach has been recently proposed to automatically learn semantically-equivalent translation rules. Because translation rules directly impact the quality and performance of the translated host codes, one of the key issues is to collect as many translation rules as possible through minimal training data set. The collected translation rules should also cover (i.e. apply to) as many guest binary instructions or code sequences as possible at the runtime. For those guest binary instructions that are not covered by the learned rules, emulation has to be used, which will incur additional runtime overhead. Prior learning-based DBT systems only achieve an average of about 69% dynamic code coverage for SPEC CINT 2006.In this paper, we propose a novel parameterization approach to take advantage of the regularity and the well-structured format in most modern ISAs. It allows us to extend the learned translation rules to include instructions or instruction sequences of similar structures or characteristics that are not covered in the training set. More translation rules can thus be harvested from the same training set. Experimental results on QEMU 4.1 show that using such a parameterization approach we can expand the learned 2,724 rules to 86,423 applicable rules for SPEC CINT 2006. Its code coverage can also be expanded from about 69.7% to about 95.5% with a 24% performance improvement compared to enhanced learning-based approach.
Jinhu Jiang, Rongchao Dong, Zhongjun Zhou, Changheng Song, Wenwen Wang 0001, Pen-Chung Yew
MICRO2