Yizhuo Rao

dblp:302/7398 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2026
0009-0000-3572-6969ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 66% Hardware accelerators and domain-specific architectures · 17% Distributed systems · 13%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
particle-in-cell simulation
2.022026
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication · HPDC 2026
Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations · EuroSys 2026
High-performance computing
scientific computing systems
2.022026
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication · HPDC 2026
Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations · EuroSys 2026
Hardware accelerators and domain-specific architectures › scientific computing accelerator › linear algebra accelerator
matrix processing units
1.322026
Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations · EuroSys 2026
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication · HPDC 2026
Distributed systems › communication optimization
communication-computation overlap
1.012026
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication · HPDC 2026
High-performance computing
performance optimization at scale
1.012026
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication · HPDC 2026
Processor architecture and microarchitecture
many-core architecture
0.312026
Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations · EuroSys 2026

Methods — techniques the papers use, named apart from their topics

outer-product formulation · 1.0matrix outer-product · 1.0co-design · 1.0asynchronous communication · 1.0
YearPublicationVenuePosition
2026 Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell Simulations
abstract
Particle-in-Cell (PIC) simulations devote most cycles to particle-grid interactions, and their fine-grained atomic updates become a severe bottleneck on traditional many-core CPUs. The evolution of CPU architectures, particularly the integration of specialized Matrix Processing Units (MPUs) designed for efficient matrix outer-product operations, presents a paradigm shift and an opportunity to alleviate these bottlenecks. Capitalizing on this architectural advancement, this work focuses on adapting the critical current deposition step in PIC simulations to this new matrix-centric computational model.
Yizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang, Guangnan Feng, Jinhui Wei, Zhiguang Chen 0001, Yutong Lu
EuroSys1
2026 POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication
abstract
Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle–grid interaction bottlenecks and particle redistribution costs. Specifically, the particle–grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-synchronous redistribution further destroys long-term data locality thereby limiting parallel efficiency. To address these inefficiencies, we present POLAR-PIC, a co-designed framework for large-scale PIC simulations that (i) reformulates Field Interpolation into an MPU-friendly outer-product form, (ii) maintains a physically ordered particle layout to preserve memory contiguity, and (iii) overlaps particle communication with Deposition to hide redistribution overhead. The evaluation on the pilot system of an Exascale supercomputer demonstrates that POLAR-PIC accelerates the entire particle-processing phase by up to 10.9 × in uniform plasma and 4.4 × in real-world laser-ion acceleration scenarios compared to the native WarpX reference pipeline on LX2. Ablation studies reveal that the speedups achieved by Interpolation and Deposition are 8.0 × and 13.2 × , respectively, and the asynchronous communication design sustains a \(99.1\%\) overlap ratio. In cross-platform comparisons, POLAR-PIC achieves \(13.2\%\) of theoretical peak efficiency on the CPU-based LS system, while WarpX reaches \(9.6\%\) on NVIDIA A800 GPUs. Notably, the scalability evaluation demonstrates that POLAR-PIC maintains \(67.5\%\) weak scaling efficiency on over 2 million cores under high-migration dynamic workloads, highlighting the importance of holistic co-design for future matrix-centric HPC systems.
Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Jinhui Wei, Languang Gao, Zhiguang Chen 0001, Yutong Lu
HPDC1
2021 Know-GNN: An Explainable Knowledge-Guided Graph Neural Network for Fraud Detection
Yizhuo Rao, Xianya Mi, Chengyuan Duan, Xiaoguang Ren, Hongliang You, Zhixian Zeng
ICONIP (5)1
2021 An Improved Ant Colony Algorithm for UAV Path Planning in Uncertain Environment
abstract
Track planning for drones has been a common problem. In order to ensure that the UAV can fly long distance in accordance with the predetermined path, it is necessary to set calibration points on the flight path to correct the sensor errors of the UAV. This problem can be abstracted into a path planning problem and solved by ant colony algorithm. Considering the uncertainty of correction failure in these calibration points. This paper presents an improved ant colony algorithm and set up the “Enhanced pheromone volatilization strategy“ to ensure that the UAV could reach the destination with the greatest possibility in this uncertain situation. We verify our algorithm on public data sets**. On data set 1, our algorithm has a 100% probability of reaching the destination, while the traditional ant colony algorithm has only a 61% probability of reaching the destination. On data set 2, our algorithm has a 56% probability of reaching the destination, while the traditional ant colony algorithm cannot find a path can reach the destination. The algorithm code*** in this paper is simple to implement, strong robustness, and can be extended to other scenarios.
Yizhuo Rao, Jianjun Cao, Zhixian Zeng, Chengyuan Duan
IJCNN1
2021 Knowledge-Guided Fraud Detection Using Semi-supervised Graph Neural Network
Yizhuo Rao, Xiaoguang Ren, Chengyuan Duan, Xianya Mi, Hongliang You, Zhixian Zeng
WISE (1)1