Yaoxu Song

dblp:339/0383 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0001-8746-0488ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 57% Storage systems · 43%
Artificial intelligence
1 paper
Graph learning · 100%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
distributed storage
1.012026
UpFuzz: Detecting Data Format Incompatibility Bugs during Distributed Storage System Upgrade · NSDI 2026
Machine learning › Graph learning
graph neural network
0.712023
uGrapher: High-Performance Graph Operator Computation via Unified Abstraction for Graph Neural Networks · ASPLOS (2) 2023
Hardware accelerators and domain-specific architectures › machine learning accelerator
graph neural network accelerator
0.712023
uGrapher: High-Performance Graph Operator Computation via Unified Abstraction for Graph Neural Networks · ASPLOS (2) 2023
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.712023
uGrapher: High-Performance Graph Operator Computation via Unified Abstraction for Graph Neural Networks · ASPLOS (2) 2023
Software testing
fuzzing
0.312026
UpFuzz: Detecting Data Format Incompatibility Bugs during Distributed Storage System Upgrade · NSDI 2026

Methods — techniques the papers use, named apart from their topics

fuzzing · 2.0unified operator abstraction · 1.3schedule strategies · 1.3adaptive parallelism · 1.3
YearPublicationVenuePosition
2026 UpFuzz: Detecting Data Format Incompatibility Bugs during Distributed Storage System Upgrade
P. C. Sruthi, Yayu Wang, Yaoxu Song, Bishal Basak Papan, Pedro Fonseca 0001, Yongle Zhang 0007
NSDI4
2023 uGrapher: High-Performance Graph Operator Computation via Unified Abstraction for Graph Neural Networks
abstract
As graph neural networks (GNNs) have achieved great success in many graph learning problems, it is of paramount importance to support their efficient execution. Different graphs and different operators present different patterns during execution. However, there is still a gap in the existing GNN acceleration research to explore adaptive parallelism. We show that existing GNN frameworks rely on handwritten static kernels, which fail to achieve the best performance across different graph operators and input graph structures. In this work, we propose uGrapher, a unified interface that achieves general high performance for different graph operators and datasets. The existing GNN frameworks can easily integrate our design for its simple and unified API. We take a principled approach that decouples a graph operator’s computation and schedule to achieve that. We first build a GNN-specific operator abstraction that incorporates the semantics of graph tensors and graph loops. We explore various schedule strategies based on the abstraction that can balance the well-established trade-off relationship between parallelism, locality, and efficiency. Our evaluation shows that uGrapher can bring up to 29.1× (3.5× on average) performance improvement over the state-of-the-art baselines on two studied NVIDIA GPUs.
Yangjie Zhou 0001, Jingwen Leng, Yaoxu Song, Shuwen Lu, Chao Li 0009, Minyi Guo, Wenting Shen, Yong Li 0045, Wei Lin 0016, Xiangwen Liu
ASPLOS (2)3
2023 AdaptGear: Accelerating GNN Training via Adaptive Subgraph-Level Kernels on GPUs
abstract
Graph neural networks (GNNs) are powerful tools for exploring and learning from graph structures and features. As such, achieving high-performance execution for GNNs becomes crucially important. Prior works have proposed to explore the sparsity (i.e., low density) in the input graph to accelerate GNNs, which uses the full-graph-level or block-level sparsity format. We show that they fail to balance the sparsity benefit and kernel execution efficiency. In this paper, we propose a novel system, referred to as AdaptGear, that addresses the challenge of optimizing GNNs performance by leveraging kernels tailored to the density characteristics at the subgraph level. Meanwhile, we also propose a method that dynamically chooses the optimal set of kernels for a given input graph. Our evaluation shows that AdaptGear can achieve a significant performance improvement, up to 6.49× (1.87× on average), over the state-of-the-art works on two mainstream NVIDIA GPUs across various datasets.
Yangjie Zhou 0001, Yaoxu Song, Jingwen Leng, Zihan Liu 0002, Weihao Cui, Zhendong Zhang 0004, Cong Guo 0003, Quan Chen 0002, Li Li 0012, Minyi Guo
CF2