Kaiyuan Rong

dblp:294/4147 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0003-7944-1211ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 100%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 57% Processor architecture and microarchitecture · 26% Performance modeling and evaluation · 17%
Network and information security
1 paper
Hardware security and side channels · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization › deep learning compiler
tensor program optimization
1.222023
Optimizing DNNs With Partially Equivalent Transformations and Automated Corrections · IEEE Trans. Computers 2023
PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections · OSDI 2021
Hardware security and side channels
microarchitectural side channel
1.012026
OCCUPY+PROBE: Cross-Privilege Branch Target Buffer Side-Channel Attacks at Instruction Granularity · NDSS 2026
Processor architecture and microarchitecture
branch prediction
0.312026
OCCUPY+PROBE: Cross-Privilege Branch Target Buffer Side-Channel Attacks at Instruction Granularity · NDSS 2026
Performance modeling and evaluation › parallel system performance
strong and weak scaling
0.212023
Critique of "A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery" by SCC Team From Tsinghua University · IEEE Trans. Parallel Distributed Syst. 2023

Methods — techniques the papers use, named apart from their topics

automated correction · 2.3instruction-granularity probing · 2.0cross-privilege attack · 2.0mutation manager · 1.3multi-linearity exploitation · 1.3program transformation · 1.0parallelization · 0.7communication overhead analysis · 0.7
YearPublicationVenuePosition
2026 OCCUPY+PROBE: Cross-Privilege Branch Target Buffer Side-Channel Attacks at Instruction Granularity
Kaiyuan Rong, Junqi Fang, Haixia Wang 0001, Dapeng Ju, Dongsheng Wang 0002
NDSS1
2023 Optimizing DNNs With Partially Equivalent Transformations and Automated Corrections
abstract
Deep neural network (DNN) applications are typically represented by tensor programs. To boost the performance of DNN computations, existing works adopt fully equivalent transformations for tensor program optimization by guaranteeing the equivalence on each element of tensors. However, as there are thousands of elements in a tensor, such optimization misses the opportunities that allow the in-equivalence of minority elements. In this work, we proposePet, the first work that introduces partially equivalent transformations to optimize tensor programs. To maintain the functional equivalence of tensor programs,Petautomatically finds and corrects the in-equivalent positions by leveraging the multi-linearity of DNN computations.Petfurther uses a mutation manager to improve search efficiency. Evaluation results show thatPetcan achieve up to 1.98$\times$and 2.20$\times$speedups on NVIDIA Tesla A100 and V100 respectively compared with existing DNN frameworks by introducing new optimization opportunities of partially equivalent transformations.
Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Feng Zhang 0007, Tuowei Wang, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Kaiyuan Rong, Yuanyong Chen
IEEE Trans. Computers10
2023 Critique of "A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery" by SCC Team From Tsinghua University
abstract
Srivastava et al. propose a parallel framework to optimize Bayesian network learning in the SC20 article entitled “A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery”. They parallelize all the phases in network constructing algorithms to achieve high performance and scalability. In this article, we reproduce the strong scaling and weak scaling experiments in that SC article. We conduct experiments on a 4-node cluster with Intel CPUs provided by the SCC committee. We further analyze the results of communication overhead. Our results show that the proposed method in that SC article scales well on the provided cluster, in accordance with the SC article.Author: Please confirm or add details for any funding or financial support for the research of this article. ?>
Juncheng Cao, Kaiyuan Rong, Mingshu Zhai, Yanyu Ren, Yuxi Zhu, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.2
2021 HyQuas: hybrid partitioner based quantum circuit simulation system on GPU
abstract
Quantum computing has shown its strong potential in solving certain important problems. Due to the intrinsic limitations of current real quantum computers, quantum circuit simulation still plays an important role in both research and development of quantum computing. GPU-based quantum circuit simulation has been explored due to GPU's high computation capability. Despite previous efforts, existing quantum circuit simulation systems usually rely on a single method to improve poor data locality caused by complex quantum entanglement. However, we observe that existing simulation methods show significantly different performance for different circuit patterns. The optimal performance cannot be obtained only with any single method.
Chen Zhang 0001, Haojie Wang 0004, Kaiyuan Rong, Jidong Zhai
ICS4
2021 PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections
Haojie Wang 0004, Jidong Zhai, Mingyu Gao 0001, Zixuan Ma, Shizhi Tang, Liyan Zheng 0001, Yuanzhi Li, Kaiyuan Rong, Yuanyong Chen
OSDI8