Cunyang Wei

dblp:337/7607 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0001-8910-4951ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Skew-aware Adaptive All-to-allv Algorithms for Dynamic Deep Learning Workloads
Cunyang Wei, Abhinav Bhatele
ICS1
2026 The Big Send-off: Scalable and Performant Collectives for Deep Learning
abstract
Collective communication is becoming increasingly important in data center and supercomputer workloads with an increase in distributed AI related jobs. However, existing libraries that provide collective support such as NCCL, RCCL, and Cray-MPICH exhibit several performance and scalability limitations on modern GPU supercomputers. To address these challenges, we introduce the Performant Collective Communication Library (PCCL), specifically targeted for distributed deep learning (DL) workloads. PCCL provides highly optimized implementations of key collectives used in distributed DL: all-gather, reduce-scatter, and all-reduce. PCCL uses a hierarchical design with learning-based adaptive selection of the best performing algorithms to scale efficiently to thousands of GPUs. It achieves substantial performance speedups over RCCL on 2048 GCDs of Frontier -- up to 168x for reduce-scatter, 33x for all-gather and 10x for all-reduce. More modest but still significant gains up to 5.7x over NCCL are observed on Perlmutter. These gains translate directly to performance improvement of production DL workloads: up to 4.9x speedup over RCCL in DeepSpeed ZeRO-3 training, and up to 2.4x speedup in DDP training.
Keshav Pradeep, Mahua Singh, Cunyang Wei, Abhinav Bhatele
IPDPS4
2026 The Case of the Elusive Application Performance on Production GPU Supercomputers
Cunyang Wei, Keshav Pradeep, Abhinav Bhatele
IPDPS1
2025 Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
abstract
Graph neural networks (GNNs) leverage the connectivity and structure of real-world graphs to learn intricate properties and relationships between nodes. Many real-world graphs exceed the memory capacity of a GPU due to their sheer size, and training GNNs on such graphs requires techniques such as mini-batch sampling to scale. The alternative approach of distributed full-graph training suffers from high communication overheads and load imbalance due to the irregular structure of graphs. We propose a three-dimensional (3D) parallel approach for full-graph training that tackles these issues and scales to billion-edge graphs. In addition, we introduce optimizations such as a double permutation scheme for load balancing, and a performance model to predict the optimal 3D configuration of our parallel implementation – Plexus. We evaluate Plexus on six different graph datasets and show scaling results on up to 2048 GPUs of Perlmutter, and 1024 GPUs of Frontier. Plexus achieves unprecedented speedups of 2.3 − 12.5 × over prior state of the art, and a reduction in time-to-solution by 5.2 − 8.7 × on Perlmutter and 7.0 − 54.2 × on Frontier.
Aditya K. Ranjan, Cunyang Wei, Abhinav Bhatele
SC3
2024 VNEC: A Vectorized Non-Empty Column Format for SpMV on CPUs
abstract
Sparse matrix-vector multiplication (SpMV) is a widely used computational kernel for many applications. The performance of existing vectorization-oriented and locality-optimized SpMV works is limited by increasing additional memory accesses to the output vector or using expensive gather operations. To address these issues, we present the Vectorized Non-Empty Column (VNEC), a novel SpMV storage format aiming to optimize locality and vectorization while alleviating the existing limitations. The VNEC chunks the sparse matrix by rows and removes the empty columns from each row block to improve input vector locality and reduce extra output vector memory access. It can also relieve the cost of expensive gather operations by padding zeros and employing less costly vector load instruction. Specifically, we design two variants of VNEC for different non-zero distributions and propose an effective heuristic selection model by introducing the Intra-Row Density (IRD) to evaluate which variant is suitable for optimizing a given matrix. Experimental results show that in a multicore environment, VNEC achieves up to 6.94× speedup (2.10× on average) against the standard MKL SpMV routine on the x86 CPU and up to 5.92× speedup (1.73× on average) over ArmPL on the ARM CPU. We emphasize that the VNEC format is practical for real-world iterative applications because of its low preprocessing overhead for format conversion.
Haipeng Jia, Lei Xu 0023, Cunyang Wei, Kun Li 0016, Xianmeng Jiang, Yunquan Zhang
IPDPS4
2024 IrGEMM: An Input-Aware Tuning Framework for Irregular GEMM on ARM and X86 CPUs
abstract
The matrix multiplication algorithm is a fundamental numerical technique in linear algebra and plays a crucial role in many scientific computing applications. Despite the high performance of mainstream basic linear algebra libraries for large-scale dense matrix multiplications, they exhibit poor performance when applied to matrix multiplication with irregular input. This paper proposes an input-aware tuning framework that accounts for application scenarios and computer architectures to provide high-performance irregular matrix multiplication on ARMv8 and X86 CPUs. The framework comprises two stages: the install-time stage and the run-time stage. The install-time stage utilizes our proposed computational template to generate high-performance kernels for general data layout and SIMD-friendly data layout. The run-time stage utilizes a tiling algorithm suitable for irregular GEMM to select the optimal kernel and link as an execution plan. Additionally, load-balanced multi-threaded optimization algorithms are defined to exploit the multi-threading capability of modern processors. Experiments demonstrate that the proposed IrGEMM framework can achieve significant performance improvements for irregular GEMM on both ARMv8 and X86 CPUs compared to other mainstream BLAS libraries.
Cunyang Wei, Haipeng Jia, Yunquan Zhang, Jianyu Yao, Chendi Li, Wenxuan Cao
IEEE Trans. Parallel Distributed Syst.1
2023 SA_TRSM: A Shape-Aware Auto-Tuning Framework for Small-Scale Irregular-Shaped TRSM
abstract
TRSM (Triangular Solve with Matrix) is an algorithm in the BLAS library for efficiently solving systems of linear equations, which is widely used in scientific computing, engineering computing, and machine learning. The traditional TRSM algorithm performs well in solving large-scale converging squareshaped matrices but is inefficient in solving small-scale irregularshaped matrices. In this paper, we propose SATRSM, a Shape-Aware auto-tuning framework that is aware of scale size and irregularity, aiming to improve performance on small-scale irregular-shaped TRSM computations. SA TRSM consists of the install-time stage and the run-time stage. In the install-time stage, we designed five components for generating high-performance kernels. In the run-time stage, we designed the Shape-Aware tiling algorithm and Plan Generator for generating an efficient execution plan. The experimental results show that the average performance of SA TRSM in this paper improves by 29.4,16.1,24.6 times, and 7.8 times on double-precision real, single-precision real, doubleprecision complex, and single-precision complex in turn, relative to the algorithms in MKL.
Rongyuan Guo, Haipeng Jia, Yunquan Zhang, Mingsen Deng, Cunyang Wei, Wenbin Chang
ICPADS5
2022 IATF: An Input-Aware Tuning Framework for Compact BLAS Based on ARMv8 CPUs
abstract
Recently the mainstream basic linear algebra libraries have delivered high performance on large scale General Matrix Multiplication(GEMM) and Triangular System Solve(TRSM). However, these libraries are still insufficient to provide sustained performance for batch operations on large groups of fixed-size small matrices on specific architectures, which are extensively used in various scientific computing applications. In this paper, we propose IATF, an input-aware tuning framework for optimizing large group of fixed-size small GEMM and TRSM to boost near-optimal performance on ARMv8 architecture. The IATF contains two stages: install-time stage and run-time stage. In the install-time stage, based on SIMD-friendly data layout, we propose computing kernel templates for high-performance GEMM and TRSM, analyze optimal kernel sizes to increase computational instruction ratio, and design kernel optimization strategies to improve kernel execution efficiency. Furthermore, an optimized data packing strategy is also presented for computing kernels to minimize the cost of memory accessing overhead. In the run-time stage, we present an input-aware tuning method to generate an efficient execution plan for large group of fixed-size small GEMM and TRSM, according to the input matrix properties. The experimental results show that IATF could achieve significant performance improvements in GEMM and TRSM compared with other mainstream BLAS libraries.
Cunyang Wei, Haipeng Jia, Yunquan Zhang, Liusha Xu
ICPP1