EDBT 2026 Demo / reviewers in the wild / expert
Yangde Feng
dblp:20/624
· DBLP profile ↗
6ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0002-0205-5561ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Scalable Hybrid Total FETI Method for Massively Parallel FEM SimulationsabstractThe Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method plays an important role in solving large-scale and complex engineering problems. This method needs to handle numerous matrix-vector multiplications. Directly calling the vendor-optimized library for general matrix-vector multiplication (gemv) on GPU leads to low performance, since it does not consider optimizations for different matrix sizes in HTFETI, i.e. different row and column sizes. In addition, state-of-the-art graph partitioning methods cannot guarantee load balancing for HTFETI, since the matrix size is determined by the length of the subdomain boundary. To solve the problems above, we first port gemv to the multi-stream pipeline scheme and develop a new batched kernel function on GPU, which brings 15%~30% throughput improvement and 37% average GFLOPs improvement, respectively. We also propose a multi-grained load-balancing scheme based on graph repartitioning and work-stealing, and the load imbalance ratio is down to 1.05~1.09 from 1.5. We have successfully applied the scalable HTFETI method to simulate the whole core assembly of China Experimental Fast Reactor (CEFR) for steady-state analysis, and the efficiencies of weak scalability and strong scalability reach 78% and 72% on 12,288 GPUs, respectively. As far as we know, this is the first time that HTFETI has been used in large-scale and high-fidelity whole core assembly simulation. Kehao Lin, Chunbao Zhou, Ningming Nie, Jue Wang 0013, Shigang Li 0002, Yangde Feng, Yangang Wang 0002, Kehan Yao, Tiechui Yao, Jian Wan 0001 |
PPoPP | 7 |
| 2023 | ANT-MOC: Scalable Neutral Particle Transport Using 3D Method of Characteristics on Multi-GPU SystemsabstractThe Method Of Characteristic (MOC) to solve the Neutron Transport Equation (NTE) is the core of full-core simulation for reactors. High resolution is enabled by discretizing the NTE through massive tracks to traverse the 3D reactor geometry. However, the 3D full-core simulation is prohibitively expensive because of the high memory consumption and the severe load imbalance. To deal with these challenges, we develop ANT-MOC1. Specifically, we build a performance model for memory footprint, computation and communication, based on which a track management strategy is proposed to overcome the resolution bottlenecks caused by limited GPU memory. Furthermore, we implement a novel multi-level load mapping strategy to ensure load balancing among nodes, GPUs, and CUs. ANT-MOC enables a 3D full-core reactor simulation with 100 billion tracks on 16,000 GPUs, with 70.69% and 89.38% parallel efficiency for strong scalability and weak scalability, respectively. Shunde Li, Zongguo Wang, Lingkun Bu, Jue Wang 0013, Zhikuang Xin, Shigang Li 0002, Yangang Wang 0002, Yangde Feng, Peng Shi 0006, Xuebin Chi |
SC | 8 |
| 2023 | Large-Scale Simulation of Structural Dynamics Computing on GPU ClustersabstractStructural dynamics simulation plays an important role in research on reactor design and complex engineering. The Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method combined with Newmark method is an efficient way to solve large-scale structural dynamics problems. However, the sparse direct solver and the load imbalance caused by inconsistent density models are two critical issues limiting the performance and the scalability of structural dynamics computing. For the former, we propose an efficient variable-size batched method to accelerate SpMV on GPUs. For the latter, we establish an online performance prediction model, based on which we then design a novel inter-cluster subdomain fine-tuning algorithm to balance the workload of HTFETI parallel computing. We are the first to achieve the high-fidelity structural dynamics simulation of China Experimental Fast Reactor core assembly with up to 53.4 billion grids. The weak and strong scalability efficiencies reach 91.77% and 86.13% on 12,800 GPUs, respectively. Yumeng Shi, Ningming Nie, Jue Wang 0013, Kehao Lin, Chunbao Zhou, Shigang Li 0002, Kehan Yao, Shunde Li, Yangde Feng, Yangang Wang 0002 |
SC | 9 |
| 2018 | Massively Scaling the Metal Microscopic Damage Simulation on Sunway TaihuLight SupercomputerabstractThe limitation of simulation scales leads to a gap between simulation results and physical phenomena. This paper reports our efforts on increasing the scalability of metal material microscopic damage simulation on the Sunway TaihuLight supercomputer. We use a multiscale modeling approach that couples Molecular Dynamics (MD) with Kinetic Monte Carlo (KMC). According to the characteristics of metal materials, we design a dedicated data structure to record the neighbor atoms for MD, which significantly reduces the memory consumption. Data compaction and double buffer are used to reduce the data transfer overhead between the main memory and the local store. We propose an on-demand communication strategy for KMC to remarkably reduce the communication overhead. We simulate 4 * 1012 atoms on 6,656,000 master+slave cores using MD with 85% parallel efficiency. Using the coupled MD-KMC approach, we simulate 3.2 * 1010 atoms in 19.2 days temporal scale on 6,240,000 master+slave cores with runtime of 8.6 hours. Shigang Li 0002, Baodong Wu, Yunquan Zhang, Xianmeng Wang, Jianjiang Li, Changjun Hu, Jue Wang 0013, Yangde Feng, Ningming Nie |
ICPP | 8 |
| 2012 | Parallel FDTD Simulation of Photonic Crystals and Thin-Film Solar CellsabstractFinite difference time domain (FDTD) method is a robust and accurate algorithm which is widely used in computational electromagnetic field and the simulation of optical phenomenon. In this paper, parallel FDTD based on overlapped domain decomposition is used to simulate the band gap of photonic crystals and the quantum efficiency of thin film solar cells. The light-trapping effect is also analyzed by parallel FDTD, it's very important to improve light absorption. Numerical result demonstrates that the accuracy and the speedup of parallel FDTD are very high for large scale problem. Xuebin Chi, Yangde Feng, Yonghua Zhao |
PDCAT | 3 |
| 2008 | Parallelization and Acceleration Scheme of Multilevel Fast Multipole MethodabstractThe iterative methods such as BiCGStab for solving electromagnetic field integer equations have a complexity of O(N2), which can be reduced to O(N logN) by multilevel fast multipole method (MLFMM). For large scale problems, MLFMM should be parallelized, and the iterative convergence can be accelerated by preconditioners such as incomplete inverse triangular factorization preconditioner. The interpolation based on spherical harmonic transform at each level of MLFMMpsilas octree can be further accelerated by FFT. Based on this acceleration scheme tested on distributed cluster, the results show this algorithm is feasible. Yangde Feng, Xuebin Chi |
PDCAT | 2 |