Yunfei Pang

dblp:341/2461 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-5386-3067ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 T-Control: An Efficient Dynamic Tensor Rematerialization System for DNN Training
Junmin Xiao, Xiaochuan Deng, Huibing Wang, Yunfei Pang, Guangming Tan
ASPLOS (2)7
2025 Hiperti: high performance system for cross-platform code generation of transformer model inference based on MLIR
Jiashu Yao, Junmin Xiao, Baokang Xie, Shilong Xu, Yunfei Pang, Yun Song, Guangming Tan
CCF Trans. High Perform. Comput.6
2024 A Coordinated Strategy for GNN Combining Computational Graph and Operator Optimizations
abstract
Graph Neural Networks (GNNs) have garnered significant interest across various domains due to their efficacy in learning from graph-structured data. In pursuit of heightened performance, numerous GNN frameworks have emerged recently. However, recent work tends to study performance optimization at the computational graph level and operator level separately, and the existing optimization techniques rely on pattern matching and manual intervention, driven by human expertise. Consequently, their performances remain sub-optimal and sensitive to input graphs and GNN models. In this work, we develop an efficient coordinated strategy named AlphaGNN, which achieves an effective combination of computational graph optimization and operator optimization. To render this coordinated optimization impactful, a rule-based computational graph optimization and a performance-driven operator optimization are proposed. The experimental results confirm that AlphaGNN achieves up to 12.39 × (2.94 × on average) performance improvement over the state-of-the-art methods on diverse GNN models.
Junmin Xiao, Zhiheng Lin, Chaoyang Shui, Yunfei Pang, Guangming Tan
ICS8
2023 Adaptive Workload-Balanced Scheduling Strategy for Global Ocean Data Assimilation on Massive GPUs
abstract
Global ocean data assimilation is a crucial technique to estimate the actual oceanic state by combining numerical model outcomes and observation data, which is widely used in climate research. Due to the imbalanced distribution of observation data in global ocean, the parallel efficiency of recent methods suffers from workload imbalance. When massive GPUs are applied for global ocean data assimilation, the workload imbalance becomes more severe, resulting in poor scalability. In this work, we propose a novel adaptive workload-balance scheduling strategy, Bassimilation, which successfully estimates the total workload prior to execution and ensures a balanced workload assignment. Further, we design a parallel dynamic programming approach to accelerate the schedule decision, and develop a factored dataflow to exploit the parallel potential of GPUs. Evaluation demonstrates that our algorithm outperforms the state-of-the-art method by up to 9.1× speedup. This work is the first to scale global ocean data assimilation to 4, 000 GPUs.
Junmin Xiao, Chaoyang Shui, Di Cai, Kangyu Wang, Yunfei Pang, Guangming Tan
SC5
2022 W-Cycle SVD: A Multilevel Algorithm for Batched SVD on GPUs
abstract
As a basic matrix factorization operation, Singular Value Decomposition (SVD) is widely used in diverse domains. In real-world applications, the computational bottleneck of matrix factorization is on small matrices, and many GPU-accelerated batched SVD algorithms have been developed recently for higher performance. However, these algorithms failed to achieve both high data locality and convergence speed, because they are size-sensitive. In this work, we propose a novel W-cycle SVD to accelerate the batched one-sided Jacobi SVD on GPUs. The W-cycle SVD, which is size-oblivious, successfully exploits the data reuse and ensures the optimal convergence speed for batched SVD. Further, we present the efficient batched kernel design, and propose a tailoring strategy based on auto-tuning to improve the batched matrix multiplication in SVDs. The evaluation demonstrates that the proposed algorithm achieves 2.6∼10.2× speedup over the state-of-the-art cuSOLVER. In a real-world data assimilation application, our algorithm achieves 2.73∼3.09× speedup compared with MAGMA.
Junmin Xiao, Yunfei Pang, Chaoyang Shui, Guangming Tan
SC2