EDBT 2026 Demo / reviewers in the wild / expert
Huang Ye
dblp:186/2177
· DBLP profile ↗
9ranked-venue papers
0as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Large-scale Phase-Field Simulations for Solid-Solid Phase Transformations involving Elastic EnergyabstractPhase-field models have been used extensively in studying microstructure evolution in alloys and have the superiority of comprehending, predicting, and optimizing microstructure-sensitive macroscopic material properties. The elastic strain energy is a vital factor in modeling crystal structure formation in solid-solid phase transformations. Conventionally, it is computed in the reciprocal space according to the famous Khachaturyan-Shatalov theory. In large-scale simulations, the full-space Fourier transform becomes extremely time-consuming. Yaqian Gao, Jian Zhang 0070, Huang Ye, Xuebin Chi |
ICPP | 3 |
| 2024 | High-Performance 3D convolution on the Latest Generation Sunway ProcessorabstractThe emergence of High-Performance Computing (HPC) and Artificial Intelligence (AI) has significantly expanded the applications of three-dimensional convolutional neural networks (3D CNNs). At the same time, the next-generation Sunway supercomputer has evidenced its superior computational capabilities in the HPC+AI domain. However, complex 3D convolution remains a primary performance limitation in many applications. The optimization of tensor-like operators on the Sunway processor is usually implemented via a multi-level blocking approach, adapting to its architecture. Although it can effectively mitigate the differences in memory access latency among different memory hierarchies, the performance of 3D convolutions is still frequently limited by the transfer bandwidth. Zhichen Feng, Yaqian Gao, Shaobo Tian, Huang Ye, Jian Zhang 0070 |
ICPP | 6 |
| 2024 | A Feature Extraction Framework for 3D Scientific Voxel Object Using SVD and Neural NetworkabstractLatest advances in computational methods and high-performance computing have enabled large-scale scientific simulations to become feasible. However, efficiently analyzing the resulting large datasets remains challenging. Typically, 3D voxel data occupies huge storage space which causes serious storage pressure. Performing real-time feature extraction during simulations could mitigate this demand. In this paper, We propose a three-decker structure framework for locating and extracting voxel object features such as pose, class, and size. The framework first utilizes a 3D CNN to localize the object’s center and size along each axis, enabling adaptive cropping of the target from the original voxel space. Then the singular value decomposition (SVD) will applied to the extracted object to preliminarily extract and refine its posture. Finally, a neural network will refine the SVD-extracted pose and predict the category and size. Finally, a neural network is utilized to infer residual pose corrections as well as category and size information for the object after preliminary pose alignment via SVD decomposition. Our method achieves accurate feature extraction for voxel objects with complex poses. More importantly, we achieve pose estimation without an initial position. By storing extracted features instead of full data, we significantly reduce storage needs. We validate our approach using the phase field simulations dataset and ModelNet40 dataset. At last, based on the proposed framework we develop an in-situ feature extraction library for running with large-scale scientific computing programs and performing real-time feature extraction on the large amount of computation data it generates. As a concrete running example, we selected the microstructure evolution program governed by the phase-field method for a common run and achieved high-accuracy feature extraction. Zhichen Feng, Yaqian Gao, Huang Ye, Jian Zhang 0070 |
IJCNN | 3 |
| 2024 | A General Parallel Framework for Material Point Method Based on the p4est LibraryabstractMaterial Point Method(MPM) is widely used to simulate large deformation processes such as material fracture, collision, and fluid structure interaction. In general, an MPM simulation evolves a dynamic process in which a large number of Lagrangian particles move on top of Eulerian grids. During the process, physical quantities such as momentum and force are interpolated frequently between the particles and their corresponding grids. High Performance Computing(HPC) has become an indispensable tool for large-scale MPM simulations, while dynamic task decomposition and load balancing are critical for efficiency.In this paper, we present a general parallel framework integrating MPM and a well-established oct-tree mesh management library, p4est, aiming to improve the efficiency of large-scale MPM simulations. Through careful design of data structure and interfaces to connect the original Grid class of MPM with the p4est mesh, a highly modular structure is realized for the framework. This design allows the users to easily add new features or optimize existing ones, thereby enhancing its flexibility and reusability. On top of this, dynamic load-balancing strategies for MPM can be realized without much effort. A recommended strategy that considers both particle and grid workload is presented. In addition, the advantage of dynamic load balancing is demonstrated and analyzed through practical simulations. The versatility and efficiency of the proposed framework are also demonstrated through concrete real-world applications including penetration and building implosion simulations with up to 1.5 billion degrees of freedom. The code achieves 87% overall parallel efficiency scaling up to 2048 processors. Shaobo Tian, Huang Ye, Jian Zhang 0070 |
ISPA | 4 |
| 2024 | POSTER: Enabling Extreme-Scale Phase Field Simulation with In-situ Feature ExtractionabstractIn this paper, we present an integrated framework composed of a highly efficient phase field simulator and an in-situ feature extraction library. This novel framework enables us to conduct extreme-scale micro-structure evolution simulations while the characteristic features of each individual grain are extracted on the fly. After systematic design and optimization on the new generation Sunway supercomputer, the code scales up to 39 million cores and achieves 582 PFlops in double precision and 637 POps in mixed precision. Zhichen Feng, Yaqian Gao, Shaobo Tian, Huang Ye, Jian Zhang 0070 |
PPoPP | 5 |
| 2022 | A Fine-grained Prefetching Scheme for DGEMM Kernels on GPU with Auto-tuning CompatibilityabstractGeneral Matrix Multiplication (GEMM) is one of the fundamental kernels for scientific and high-performance computing. When optimizing the performance of GEMM on GPU, the matrix is usually partitioned into a hierarchy of tiles to fit the thread hierarchy. In practice, the thread-level parallelism is affected not only by the tiling scheme but also by the resources that each tile consumes, such as registers and local data share memory. This paper presents a fine-grained prefetching scheme that improves the thread-level parallelism by balancing the usage of such resources. The gain and loss on instruction and thread level parallelism are analyzed and a mathematical model is developed to estimate the overall performance gain. Moreover, the proposed scheme is integrated into the open-source tool Tensile to automatically generate assembly and tune a collection of kernels to maximize the performance of DGEMM for a family of problem sizes. Experiments show about 1.10X performance speedup on a wide range of matrix sizes for both single and batched matrix-matrix multiplication. Huang Ye, Shaobo Tian, Jian Zhang 0070 |
IPDPS | 2 |
| 2021 | Redesigning Peridigm on SIMT Accelerators for High-performance Peridynamics SimulationsabstractPeridigm is one of the most frequently utilized Peridynamics (PD) simulation software for problems involving discontinuity, such as cracks and fragmentation. However, performing long-term and large-scale simulations is very time-consuming for Peridigm. To enhance the performance and scalability of Peridigm, we port and optimize Peridigm on the SIMT accelerators. Challenges are imposed on efficient Peridigm on the SIMT architecture by the complex calculations and massive memory access of PD simulations. In this study, a series of strategies and techniques are proposed to optimize the performance of Peridigm. We first adjust the algorithms of bond-based calculations to eliminate the data conflicts with minimized overhead in order to achieve parallel Peridigm on accelerators. Furthermore, we propose thread grouping and collaborative memory access strategies to decrease the overhead of data fetch from device memory. To improve the efficiency of calculations, we also refine the calculation instructions. Finally, we offer a transmission-computation overlapping strategy for reducing the overhead brought by the data transmissions and improving the scalability. The optimized Peridigm on 4 Nvidia Tesla V100 GPUs accelerates the basic parallel Peridigm on 4 V100 GPUs 10.24 times. Compared to the original Peridigm run on 8 Intel Xeon Gold 6248 CPUs (160 cores, 320 threads) and the optimized PD application run on 4 SW26010 processors (1,040 cores), our work on 4 V100 GPUs accelerates the simulation 9 times and 4 times respectively. As for large-scale simulations, because we don't have enough V100 GPUs, we run our work on noncommercial SIMT accelerators which have similar performance to the V100 of the PCIe version, with the example scales from 282,000 points to 36,096,000 points and the number of accelerators scales from 4 to 512, near-linear scalability is observed and the performance ultimately reaching 825.72 TFLOPS with 98.81% parallel efficiency Huang Ye, Jian Zhang 0070 |
IPDPS | 2 |
| 2020 | Large-scale Simulations of Peridynamics on Sunway Taihulight SupercomputerabstractPeridynamics (PD) methods are good at describing solid mechanical behaviours and have the superiority on simulating the discontinuous problems. They can be applied to many fields, such as materials science, human health, and industrial manufacturing, etc., which motivates us to provide their efficient numerical simulations on the Sunway TaihuLight supercomputer. However, massive and complex calculations of PD simulations and the characteristics of Sunway TaihuLight bring challenges to efficient parallel PD simulations. In this paper, we present a series of performance optimization techniques to perform a large-scale parallel PD simulation application on Sunway TaihuLight. We first design the data grouping and SPM-based caching to increase the bandwidth of data transmission and reduce the time of the main memory access. Further, we design and implement vectorization and instruction-level optimization for PD applications to improve computational performance. Finally, we offer the overlapping strategies of data transmission and computation so that data transmission can be covered by computation. Our work in a core group improves the performance of the serial version on the SW26010 processor by 181 times. Compared to the serial and single-CPU Peridigm-based simulations on Intel Xeon E5-2680 V3, our work gets a speedup of 60 times and 6 times, respectively. Near linear scalability is also obtained. When testing the weak scaling, the simulation of a 296,222,720-point example achieves 1.14 PFLOPS with 8192 (532,480 cores) processes. When testing the strong scaling, 90% parallel efficiency is observed as the number of processes increases 64 times to 4096 processes. Huang Ye, Jian Zhang 0070 |
ICPP | 2 |
| 2016 | A distributed load balancing algorithm for climate big data processing over a multi-core CPU clusterabstractSummary Load imbalance is a common problem to be tackled urgently in large scale data‐driven simulation systems or data intensive computing. According to the coupler, the Chinese Academy of Sciences‐Earth System Model (CAS‐ESM) implements one‐way nesting of the Institute of Atmospheric Physics of Chinese Academy of Sciences Atmospheric General Circulation Model version 4.0 (IAP AGCM4.0) and Weather Research and Forecasting model (WRF). The METGRID (meteorological grid) and REAL program modules in the WRF are used to process meteorological data. In the CAS‐ESM, the load of the METGRID module is seriously unbalanced on many CPU cores. The load imbalance has a serious impact on the processing speed of meteorological data, so this study designs an optimization algorithm to solve the problem. Numerical experiments show that compared to before optimization, the optimization algorithm can solve the load imbalance of the METGRID, and the computation speed of the METGRID and REAL modules after optimization on 64 CPU cores is about 7.2 times faster than before. Meanwhile, the whole computation speed of the CAS‐ESM can improve by 217.53%. In addition, results indicate that they also can reach to a similar speedup on different numbers of CPU cores. Copyright © 2016 John Wiley & Sons, Ltd. Jinrong Jiang, Huang Ye, Juanxiong He |
Concurr. Comput. Pract. Exp. | 3 |