EDBT 2026 Demo / reviewers in the wild / expert
Ningming Nie
dblp:191/1617
· DBLP profile ↗
7ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-6254-0147ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Full-Core Fluid-Structure-Interaction Simulation of Nuclear Reactor on CPU+GPU Hybrid ClustersabstractNuclear reactor FSI simulation faces two key challenges: "Mapping wall" bottleneck in data transfer across non-matching mesh coupling interfaces; Low hardware utilization from multi-physics solvers’ heterogeneous core tasks (compute- vs. memory-intensive). Therefore, an innovative FSI framework integrating two strategies is proposed: Scalable radial basis function mapping—restructuring the global problem into massive independent subproblems via task partitioning, preallocation, and multi-granularity load balancing to eliminate communication overhead; Dependency-aware multi-stream optimization—deeply overlapping heterogeneous solver tasks to maximize hardware utilization. It first achieves parameter transfer across ∼90,000 non-matching coupling interfaces in China Experimental Fast Reactor, with 86.36% strong scaling and 94.01% weak scaling. The combined optimizations yield ∼60% performance gain, increase strong scaling by over 20 percentage points, and achieve high weak scaling of ∼97%. Moreover, the FSI results align well with publicly available data, verifying its correctness. Xue Miao, Jue Wang 0013, Qida Lin, Shufei Zhang, Rongqiang Cao, Chunbao Zhou, Ningming Nie, He Bai 0005, Yangang Wang 0002 |
HPDC | 7 |
| 2025 | MCloudNet: An Ultra-Short-Term Photovoltaic Power Forecasting Framework With Multi-Layer Cloud CoverageabstractOver 4.15 million low-income households across nearly 60,000 villages in China benefit from photovoltaic (PV) poverty alleviation power stations. However, weak infrastructure and limited capabilities make these systems vulnerable to fluctuations. One of the United Nations' Sustainable Development Goals (SDG 7) seeks to ensure access to affordable and reliable energy for all, especially in underdeveloped regions. This paper proposes MCloudNet, a multi-modal framework designed to improve ultra-short-term PV prediction in data-scarce, cloud-dynamic environments. MCloudNet explicitly models multi-layer cloud structures from satellite imagery and fuses them with time-series meteorological data to enhance prediction accuracy and interpretability. A province-level dispatch system with MCloudNet has been deployed in Hebei, supporting scheduling across rural PV stations. Experiments conducted in counties such as Shexian and Luxi highlight the framework's effectiveness for use in underdeveloped micro-grids. Operational results show that the system has reduced over 60 million kWh of solar curtailment and generated 24 million CNY in economic value, benefiting approximately 50,000 rural households. By minimizing power fluctuations and improving rural energy scheduling, MCloudNet supports essential services such as lighting, medical facilities, and communications. The source code is available at: https://github.com/AI4SClab/MCloudNet. Meng Wan, Yuxuan Bi, Jue Wang 0013, Rongqiang Cao, Jiaxiang Wang 0002, Peng Shi 0006, Ningming Nie, Yangang Wang 0002 |
IJCAI | 9 |
| 2025 | MISA-AKMC : Achieve Kinetic Monte Carlo Simulation of 20 Quadrillion Atoms on GPU ClustersabstractThe Atomic Kinetic Monte Carlo (AKMC) method provides insights into the macroscopic behavior of materials through atomistic-level simulations and finds broad applications in materials science innovation. Improving simulation scale and performance remains a consistent focus in the development of parallel AKMC software. We port the AKMC software to GPU clusters. To alleviate the memory pressure in large-scale complex system simulations, we redesign the data layout and propose the Lattice Data Compression and Vacancy Data Decompression algorithms. Additionally, We propose a multi-level pipeline scheme combined with an on-demand communication forwarding and merging strategy to reduce data transfer and communication overhead. Compared to state-of-the-art KMC software, MISA-AKMC achieves a 10.41-fold improvement in computational throughput and a 52.07-fold expansion in simulation scale. We implement the first true micrometer-scale AKMC simulation involving 20 quadrillion atoms on GPU clusters. MISA-AKMC achieves 96.03% parallel efficiency in weak scaling and 85.29% in strong scaling on 16,000 GPUs. Shunde Li, Ningming Nie, Jue Wang 0013, He Bai 0005, Genshen Chu, Xinfu He, Yangang Wang 0002, Changjun Hu, Xuebin Chi |
SC | 3 |
| 2023 | A Scalable Hybrid Total FETI Method for Massively Parallel FEM SimulationsabstractThe Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method plays an important role in solving large-scale and complex engineering problems. This method needs to handle numerous matrix-vector multiplications. Directly calling the vendor-optimized library for general matrix-vector multiplication (gemv) on GPU leads to low performance, since it does not consider optimizations for different matrix sizes in HTFETI, i.e. different row and column sizes. In addition, state-of-the-art graph partitioning methods cannot guarantee load balancing for HTFETI, since the matrix size is determined by the length of the subdomain boundary. To solve the problems above, we first port gemv to the multi-stream pipeline scheme and develop a new batched kernel function on GPU, which brings 15%~30% throughput improvement and 37% average GFLOPs improvement, respectively. We also propose a multi-grained load-balancing scheme based on graph repartitioning and work-stealing, and the load imbalance ratio is down to 1.05~1.09 from 1.5. We have successfully applied the scalable HTFETI method to simulate the whole core assembly of China Experimental Fast Reactor (CEFR) for steady-state analysis, and the efficiencies of weak scalability and strong scalability reach 78% and 72% on 12,288 GPUs, respectively. As far as we know, this is the first time that HTFETI has been used in large-scale and high-fidelity whole core assembly simulation. Kehao Lin, Chunbao Zhou, Ningming Nie, Jue Wang 0013, Shigang Li 0002, Yangde Feng, Yangang Wang 0002, Kehan Yao, Tiechui Yao, Jian Wan 0001 |
PPoPP | 4 |
| 2023 | Large-Scale Simulation of Structural Dynamics Computing on GPU ClustersabstractStructural dynamics simulation plays an important role in research on reactor design and complex engineering. The Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method combined with Newmark method is an efficient way to solve large-scale structural dynamics problems. However, the sparse direct solver and the load imbalance caused by inconsistent density models are two critical issues limiting the performance and the scalability of structural dynamics computing. For the former, we propose an efficient variable-size batched method to accelerate SpMV on GPUs. For the latter, we establish an online performance prediction model, based on which we then design a novel inter-cluster subdomain fine-tuning algorithm to balance the workload of HTFETI parallel computing. We are the first to achieve the high-fidelity structural dynamics simulation of China Experimental Fast Reactor core assembly with up to 53.4 billion grids. The weak and strong scalability efficiencies reach 91.77% and 86.13% on 12,800 GPUs, respectively. Yumeng Shi, Ningming Nie, Jue Wang 0013, Kehao Lin, Chunbao Zhou, Shigang Li 0002, Kehan Yao, Shunde Li, Yangde Feng, Yangang Wang 0002 |
SC | 2 |
| 2022 | A parallel ETD algorithm for large-scale rate theory simulation
Jianjiang Li, Baixue Ji, Xinfu He, Ningming Nie |
J. Supercomput. | 7 |
| 2018 | Massively Scaling the Metal Microscopic Damage Simulation on Sunway TaihuLight SupercomputerabstractThe limitation of simulation scales leads to a gap between simulation results and physical phenomena. This paper reports our efforts on increasing the scalability of metal material microscopic damage simulation on the Sunway TaihuLight supercomputer. We use a multiscale modeling approach that couples Molecular Dynamics (MD) with Kinetic Monte Carlo (KMC). According to the characteristics of metal materials, we design a dedicated data structure to record the neighbor atoms for MD, which significantly reduces the memory consumption. Data compaction and double buffer are used to reduce the data transfer overhead between the main memory and the local store. We propose an on-demand communication strategy for KMC to remarkably reduce the communication overhead. We simulate 4 * 1012 atoms on 6,656,000 master+slave cores using MD with 85% parallel efficiency. Using the coupled MD-KMC approach, we simulate 3.2 * 1010 atoms in 19.2 days temporal scale on 6,240,000 master+slave cores with runtime of 8.6 hours. Shigang Li 0002, Baodong Wu, Yunquan Zhang, Xianmeng Wang, Jianjiang Li, Changjun Hu, Jue Wang 0013, Yangde Feng, Ningming Nie |
ICPP | 9 |