EDBT 2026 Demo / reviewers in the wild / expert
Shaobo Tian
dblp:191/0523
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2024
0009-0009-6645-0438ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | High-Performance 3D convolution on the Latest Generation Sunway ProcessorabstractThe emergence of High-Performance Computing (HPC) and Artificial Intelligence (AI) has significantly expanded the applications of three-dimensional convolutional neural networks (3D CNNs). At the same time, the next-generation Sunway supercomputer has evidenced its superior computational capabilities in the HPC+AI domain. However, complex 3D convolution remains a primary performance limitation in many applications. The optimization of tensor-like operators on the Sunway processor is usually implemented via a multi-level blocking approach, adapting to its architecture. Although it can effectively mitigate the differences in memory access latency among different memory hierarchies, the performance of 3D convolutions is still frequently limited by the transfer bandwidth. Zhichen Feng, Yaqian Gao, Shaobo Tian, Huang Ye, Jian Zhang 0070 |
ICPP | 4 |
| 2024 | A General Parallel Framework for Material Point Method Based on the p4est LibraryabstractMaterial Point Method(MPM) is widely used to simulate large deformation processes such as material fracture, collision, and fluid structure interaction. In general, an MPM simulation evolves a dynamic process in which a large number of Lagrangian particles move on top of Eulerian grids. During the process, physical quantities such as momentum and force are interpolated frequently between the particles and their corresponding grids. High Performance Computing(HPC) has become an indispensable tool for large-scale MPM simulations, while dynamic task decomposition and load balancing are critical for efficiency.In this paper, we present a general parallel framework integrating MPM and a well-established oct-tree mesh management library, p4est, aiming to improve the efficiency of large-scale MPM simulations. Through careful design of data structure and interfaces to connect the original Grid class of MPM with the p4est mesh, a highly modular structure is realized for the framework. This design allows the users to easily add new features or optimize existing ones, thereby enhancing its flexibility and reusability. On top of this, dynamic load-balancing strategies for MPM can be realized without much effort. A recommended strategy that considers both particle and grid workload is presented. In addition, the advantage of dynamic load balancing is demonstrated and analyzed through practical simulations. The versatility and efficiency of the proposed framework are also demonstrated through concrete real-world applications including penetration and building implosion simulations with up to 1.5 billion degrees of freedom. The code achieves 87% overall parallel efficiency scaling up to 2048 processors. Shaobo Tian, Huang Ye, Jian Zhang 0070 |
ISPA | 1 |
| 2024 | POSTER: Enabling Extreme-Scale Phase Field Simulation with In-situ Feature ExtractionabstractIn this paper, we present an integrated framework composed of a highly efficient phase field simulator and an in-situ feature extraction library. This novel framework enables us to conduct extreme-scale micro-structure evolution simulations while the characteristic features of each individual grain are extracted on the fly. After systematic design and optimization on the new generation Sunway supercomputer, the code scales up to 39 million cores and achieves 582 PFlops in double precision and 637 POps in mixed precision. Zhichen Feng, Yaqian Gao, Shaobo Tian, Huang Ye, Jian Zhang 0070 |
PPoPP | 4 |
| 2022 | A Fine-grained Prefetching Scheme for DGEMM Kernels on GPU with Auto-tuning CompatibilityabstractGeneral Matrix Multiplication (GEMM) is one of the fundamental kernels for scientific and high-performance computing. When optimizing the performance of GEMM on GPU, the matrix is usually partitioned into a hierarchy of tiles to fit the thread hierarchy. In practice, the thread-level parallelism is affected not only by the tiling scheme but also by the resources that each tile consumes, such as registers and local data share memory. This paper presents a fine-grained prefetching scheme that improves the thread-level parallelism by balancing the usage of such resources. The gain and loss on instruction and thread level parallelism are analyzed and a mathematical model is developed to estimate the overall performance gain. Moreover, the proposed scheme is integrated into the open-source tool Tensile to automatically generate assembly and tune a collection of kernels to maximize the performance of DGEMM for a family of problem sizes. Experiments show about 1.10X performance speedup on a wide range of matrix sizes for both single and batched matrix-matrix multiplication. Huang Ye, Shaobo Tian, Jian Zhang 0070 |
IPDPS | 3 |