Xiuzhu Sha

dblp:321/8657 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 38% GPUs and heterogeneous computing · 30% High-performance computing · 28%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › GPU communication
GPU cluster communication
2.132026
HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters · NSDI 2026
Canvas: Scalable Collective Communication Scheduling for Large-Scale GPU Clusters · ICNP 2025
TuCCL: Tailored and Unified Configuration Optimizations for High-Performance Collective Communication Library · ICNP 2025
Distributed systems › distributed scheduling
collective communication scheduling
1.922026
HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters · NSDI 2026
Canvas: Scalable Collective Communication Scheduling for Large-Scale GPU Clusters · ICNP 2025
High-performance computing
collective communication
1.722025
TuCCL: Tailored and Unified Configuration Optimizations for High-Performance Collective Communication Library · ICNP 2025
Canvas: Scalable Collective Communication Scheduling for Large-Scale GPU Clusters · ICNP 2025
Distributed systems › distributed machine learning
distributed training systems
0.912025
TuCCL: Tailored and Unified Configuration Optimizations for High-Performance Collective Communication Library · ICNP 2025
Electronic design automation › design optimization
topology optimization
0.312026
HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters · NSDI 2026
High-performance computing › large-scale training
large-scale distributed training
0.312025
Canvas: Scalable Collective Communication Scheduling for Large-Scale GPU Clusters · ICNP 2025

Methods — techniques the papers use, named apart from their topics

scheduling algorithm · 1.0optimization · 1.0topology-aware sketch generation · 0.9pipeline scheduling · 0.9multi-phase resource-aware configuration optimization · 0.9hierarchical synthesis · 0.9hierarchical configuration optimization · 0.9collective decomposition · 0.9
YearPublicationVenuePosition
2026 DistSRE: Synthesizing Runtime Executions with Automated Co-Optimization for Distributed LLM Training
Xiuzhu Sha, Chenyang Hei, Fuliang Li, Chengxi Gao, Rongfei Zeng, Xingwei Wang 0001
IWQoS1
2026 HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters
Chenyang Hei, Jiamin Cao, Chengxi Gao, Xiuzhu Sha, Tongrui Liu, Dengke Zhang, Ennan Zhai, Xingwei Wang 0001
NSDI5
2025 Canvas: Scalable Collective Communication Scheduling for Large-Scale GPU Clusters
abstract
State-of-the-art deep learning models rely on large GPU clusters and various parallelism strategies, which in turn depend on collective communication (CC) operators to synchronize data. While vendor libraries (e.g., NCCL, RCCL) provide standard CC algorithms, they often suffer from bandwidth bottlenecks in imbalanced topologies. Recent synthesis-based methods improve performance but face three key limitations: poor scalability due to the combinatorial explosion of scheduling space, lack of support for multistage execution, and suboptimal communication throughput. We propose Canvas, a scalable and near-optimal CC scheduling framework that addresses these challenges. Canvas introduces: (1) Hierarchical synthesis to decompose the global scheduling problem into tractable subproblems for scalability. (2) Collective decomposition to enable structured, multi-stage algorithm generation. (3) Cross-micro-batch pipeline scheduling to parallelize communication across micro-batches and maximize link utilization. Evaluations show that Canvas achieves up to 1.98× bandwidth speedup over TACCL and 3.56× over TE-CCL, and synthesizes algorithms for 512-GPU topologies within 1.77 hours, whereas TACCL fails to produce results within 24 hours.
Chenyang Hei, Fuliang Li, Chengxi Gao, Tongrui Liu, Xiuzhu Sha, Xingwei Wang 0001
ICNP6
2025 TuCCL: Tailored and Unified Configuration Optimizations for High-Performance Collective Communication Library
abstract
Modern distributed training systems face escalating communication bottlenecks as GPU clusters scale to accommodate large models. While collective communication libraries and automated synthesizers address algorithmic efficiency, they suffer from three critical limitations including labor-intensive manual intervention requirement, overreliance on predefined input optimization, and suboptimal isolated configuration optimization. To solve these problems, we present TuCCL, a systematic framework that co-optimizes communication algorithms and runtime parameters through three innovations: Topology-Aware Sketch Generation that automatically produces high-performance primitives, Hierarchical Configuration Optimization modeling nonlinear parameter-performance relationships, and Multi-phase Resource-Aware Configuration Optimization enabling joint configuration tuning with adaptive search space pruning. Evaluations demonstrate TuCCL’s superiority over state-of-the-art systems with 1.75x–11.49x bandwidth improvements for AllGather/AllReduce on NVIDIA V100/A100 clusters, 90.2% faster configuration search than grid methods, and 1.22x-2.52x end-to-end training speedups across diverse model scales.
Chenyang Hei, Fuliang Li, Tongrui Liu, Chengxi Gao, Xiuzhu Sha, Xingwei Wang 0001
ICNP6