EDBT 2026 Demo / reviewers in the wild / expert
Chaoyang Shui
dblp:259/5445
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2024
0000-0003-4008-0036ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Coordinated Strategy for GNN Combining Computational Graph and Operator OptimizationsabstractGraph Neural Networks (GNNs) have garnered significant interest across various domains due to their efficacy in learning from graph-structured data. In pursuit of heightened performance, numerous GNN frameworks have emerged recently. However, recent work tends to study performance optimization at the computational graph level and operator level separately, and the existing optimization techniques rely on pattern matching and manual intervention, driven by human expertise. Consequently, their performances remain sub-optimal and sensitive to input graphs and GNN models. In this work, we develop an efficient coordinated strategy named AlphaGNN, which achieves an effective combination of computational graph optimization and operator optimization. To render this coordinated optimization impactful, a rule-based computational graph optimization and a performance-driven operator optimization are proposed. The experimental results confirm that AlphaGNN achieves up to 12.39 × (2.94 × on average) performance improvement over the state-of-the-art methods on diverse GNN models. Junmin Xiao, Zhiheng Lin, Chaoyang Shui, Yunfei Pang, Guangming Tan |
ICS | 5 |
| 2024 | Exploiting Fine-Grained Redundancy in Set-Centric Graph Pattern MiningabstractGraph Pattern Mining (GPM) applications are memory intensive as they require a tremendous amount of edge checks. In recent years, the "set-centric" abstraction has gained attention for its powerful expressive abilities. By leveraging relational algebra, they optimized algorithms with methods like matching orders, early termination, automorphism-breaking, and result reuse to reduce redundancy. However, these approaches primarily address coarse-grained redundancy from exactly the same set formulas, neglecting that the data graph's inherent locality may lead to fine-grained duplicated edge checks. In fact, even unrelated set operations may check the same pair of vertices. This paper introduces the set union operation to the set-centric abstraction to fuse duplicated edge checks into one. It maintains the expressive power of relational algebra and previous optimizations while effectively avoids fine-grained redundancy in GPM tasks. Compared to state-of-the-art methods, our method achieves significant speedup on a V100 GPU cluster, demonstrating up to 305 × faster performance than the state-of-the-art GPM system G2Miner. Zhiheng Lin, Chaoyang Shui, Junmin Xiao, Guangming Tan |
PPoPP | 3 |
| 2023 | GraphPar: Efficient Workload-Aware Subgraph Matching System on Multiple GPUsabstractSubgraph matching (SM) has witnessed tremendous progress in recent years, enabling a broad spectrum of big data applications. SM applications are extremely computeintensive since they require tremendous set operations, i.e., enumerating all the possible vertex pairs and counting the common neighbor of each pair. GPU is potentially promising hardware to accelerate SM applications due to its massive parallelism. However, SM applications achieve low efficiency and often fail to deliver high performance in multi-GPU systems owing to irregular edge distribution which exhausts the computing power and aggravates the load-imbalance problems. Although many existing frameworks have proffer numerous methods at high-level to improve the efficiency of GPU-based SM, e.g., assign matching order, early termination, and automorphismbreaking, the low-level issues on GPU architecture and system, e.g., thread mapping, graph partitions are not well addressed. In this work, we develop GraphPar, an efficient SM system targeting multi-GPUs. GraphPar proposes an effective workload- aware scheduling and an efficient set operation designing, which could successfully reduce the stragglers and significantly accelerate SM. Experiments on a V100 GPU cluster show that GraphPar is up to 4.21 × faster than the state-of-the-art GPU-based GPM system G2Miner. Junmin Xiao, Zhiheng Lin, Chaoyang Shui, Guangming Tan |
ICPADS | 5 |
| 2023 | Adaptive Workload-Balanced Scheduling Strategy for Global Ocean Data Assimilation on Massive GPUsabstractGlobal ocean data assimilation is a crucial technique to estimate the actual oceanic state by combining numerical model outcomes and observation data, which is widely used in climate research. Due to the imbalanced distribution of observation data in global ocean, the parallel efficiency of recent methods suffers from workload imbalance. When massive GPUs are applied for global ocean data assimilation, the workload imbalance becomes more severe, resulting in poor scalability. In this work, we propose a novel adaptive workload-balance scheduling strategy, Bassimilation, which successfully estimates the total workload prior to execution and ensures a balanced workload assignment. Further, we design a parallel dynamic programming approach to accelerate the schedule decision, and develop a factored dataflow to exploit the parallel potential of GPUs. Evaluation demonstrates that our algorithm outperforms the state-of-the-art method by up to 9.1× speedup. This work is the first to scale global ocean data assimilation to 4, 000 GPUs. Junmin Xiao, Chaoyang Shui, Di Cai, Kangyu Wang, Yunfei Pang, Guangming Tan |
SC | 2 |
| 2022 | W-Cycle SVD: A Multilevel Algorithm for Batched SVD on GPUsabstractAs a basic matrix factorization operation, Singular Value Decomposition (SVD) is widely used in diverse domains. In real-world applications, the computational bottleneck of matrix factorization is on small matrices, and many GPU-accelerated batched SVD algorithms have been developed recently for higher performance. However, these algorithms failed to achieve both high data locality and convergence speed, because they are size-sensitive. In this work, we propose a novel W-cycle SVD to accelerate the batched one-sided Jacobi SVD on GPUs. The W-cycle SVD, which is size-oblivious, successfully exploits the data reuse and ensures the optimal convergence speed for batched SVD. Further, we present the efficient batched kernel design, and propose a tailoring strategy based on auto-tuning to improve the batched matrix multiplication in SVDs. The evaluation demonstrates that the proposed algorithm achieves 2.6∼10.2× speedup over the state-of-the-art cuSOLVER. In a real-world data assimilation application, our algorithm achieves 2.73∼3.09× speedup compared with MAGMA. Junmin Xiao, Yunfei Pang, Chaoyang Shui, Guangming Tan |
SC | 4 |
| 2021 | Optimizing the LINPACK Algorithm for Large-Scale PCIe-Based CPU-GPU Heterogeneous SystemsabstractThere is a widening gap between GPU and other components (CPU, PCIe bus and communication network) in heterogeneous parallel system. The gap forces us to orchestrate cooperative execution among these components much more carefully than ever before. By taking the LINPACK benchmark as a case study, this article proposes a fine-grained pipelining algorithm on large-scale CPU-GPU heterogeneous cluster systems. First, we build an algorithmic model that reveals a new approach to GPU-centric and fine-grained pipelining algorithm design. Then, we present four model-driven pipelining algorithms that incrementally squeeze bubbles in the pipeline so that it is occupied by more useful floating-point calculations. The algorithms are implemented on both the AMD and NVIDIA GPU platforms. The finally optimized LINPACK program achieves 107 PFlops on 25, 600 GPUs (70 percent floating-point efficiency). Several insights have been drawn to suggest tradeoff of algorithm design, programming support, and architecture design. Guangming Tan, Chaoyang Shui, Yinshan Wang, Xianzhi Yu, Yujin Yan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Revisiting linpack algorithm on large-scale CPU-GPU heterogeneous systemsabstractAs the widening gap between GPU computing capability and other components (CPU, PCIe bus and communication network), it's increasingly challenging to design high performance parallel algorithms for large CPU-GPU heterogeneous systems. There are mainly two reasons. Firstly, simply offloading the kernel library to GPU incurs large volume data transfer through low-speed PCIe bus. Secondly, communication overheads through network severely affects scalability. To solve the above issues, we advocate a paradigm shift to CPU-centric and fine-grained pipelining algorithm design. By taking Linpack benchmark as a case study, the new algorithm design paradigm shows its effectiveness. Our optimized Linpack program achieves 63.79PFlops on 16384 GPUs. Its floating-point efficiency outperforms the NVIDIA proprietary counterparts by 5% on average. Chaoyang Shui, Xianzhi Yu, Yujin Yan, Yinshan Wang, Guangming Tan |
PPoPP | 1 |