Mingfan Li

dblp:243/2130 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-1079-3126ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
abstract
The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth. While these systems provide unified global address spaces and high-bandwidth fabrics, their full potential for sparse MoE communication is hindered by three fundamental bottlenecks: (1) Strict execution serialization imposed by coarse-grained Bulk Synchronous Parallel (BSP) orchestration of interdependent communication phases; (2) Prohibitive synchronization overhead that fails to scale alongside high interconnect bandwidth; and (3) Severe load imbalance resulting from distance-agnostic scheduling of irregular token traffic. To eliminate these bottlenecks, we introduce UBEP (Unified-Bus Expert Parallelism), a production-ready communication library that rethinks MoE's All-to-All primitives for modern superpod architectures. Through large-scale experiments, UBEP reduces All-to-All latency by up to 52.4% and MoE inference Time Per Output Token (TPOT) by up to 11.1%.
Chang Liu 0001, Si Shen, Jiaqi Zheng 0001, Mingfan Li, Yuyang Yang, Guanhua Li, Yuquan Zhang, Zhongzhe Hu, Qihang Duan, Wenkai Ling, Baochuan Yang, Xianzhi Yu, Guihai Chen
SIGCOMM5
2025 29-Billion Atoms Molecular Dynamics Simulation With Ab Initio Accuracy on 35 Million Cores of New Sunway Supercomputer
abstract
Physical phenomena such as bond breaking and phase transitions require molecular dynamics (MD) withab initioaccuracy, involving up to billions of atoms and over nanosecond timescales. Previous state-of-the-art work has demonstrated that neural network molecular dynamics (NNMD) like deep potential molecular dynamics (DeePMD), can successfully extend the temporal and spatial scales of MD withab initioaccuracy on both ARM and GPU platforms. However, the DeePMD-kit package is currently unable to fully exploit the computational potential of the new Sunway supercomputer due to its unique many-core architecture, memory hierarchy, and low precision capability. In this paper, we re-design the DeePMD-kit to harness the massive computing power of the new Sunway, enabling the MD with over ten billion atoms. We first design a large-scale parallelization scheme to exploit the massive parallelism of the new Sunway. Then we devise specialized optimizations for the time-consuming operators. Finally, we design a novel mixed precision method for DeePMD-kit customized operators to leverage the low precision computing power of the new Sunway. The optimized DeePMD-kit achieves 67.6 / 56.5$\boldsymbol{\times}$speedup for water / copper systems on the new Sunway. Meanwhile, it can perform 29 billion atoms simulation for the water system on 35 million cores (i.e., 90,000 computing nodes, around 84% of the whole supercomputer) with a peak performance of 57.1 PFLOPs, which is 7.9$\boldsymbol{\times}$bigger and 1.2$\boldsymbol{\times}$faster than state-of-the-art results. This paves the way for investigating more realistic scenarios, such as studying the mechanical properties of metals, semiconductor devices, batteries, and other materials and physical systems.
Xun Wang 0010, Xiangyu Meng 0005, Zhuoqiang Guo, Mingzhen Li 0001, Mingfan Li, Ninghui Sun, Guangming Tan, Weile Jia
IEEE Trans. Computers6
2022 AI for Quantum Mechanics: High Performance Quantum Many-Body Simulations via Deep Learning
abstract
Solving quantum many-body problems is one of the most fascinating research fields in condensed matter physics. An efficient numerical method is crucial to understand the mechanism of novel physics, such as the high Tc superconductivity, as one has to find the optimal solution in the exponentially large Hilbert space. The development of Artificial Intelligence (AI) provides a unique opportunity to solve the quantum many-body problems, but there is still a large gap from the goal. In this work, we present a novel computational framework, and adapt it to the Sunway supercomputer. With highly efficient scalability up to 40 million heterogeneous cores, we can drastically increase the number of variational parameters, which greatly improves the accuracy of the solutions. The investigations of the spin-1/2 J1-J2 model and the t-J model achieve unprecedented accuracy and time-to-solution far beyond the previous state of the art.
Xuncheng Zhao, Mingfan Li, Junshi Chen 0003, Meijia Zhao, Hong An, Lixin He
SC2
2022 Bridging the Gap between Deep Learning and Frustrated Quantum Spin System for Extreme-Scale Simulations on New Generation of Sunway Supercomputer
abstract
Efficient numerical methods are promising tools for delivering unique insights into the fascinating properties of physics, such as the highly frustrated quantum many-body systems. However, the computational complexity of obtaining the wave functions for accurately describing the quantum states increases exponentially with respect to particle number. Here we present a novel convolutional neural network (CNN) for simulating the two-dimensional highly frustrated spin-$1/2$$J_1-J_2$Heisenberg model, meanwhile the simulation is performed at an extreme scale system with low cost and high scalability. By ingenious employment of transfer learning and CNN’s translational invariance, we successfully investigate the quantum system with the lattice size up to$24\times 24$, within 30 million cores of the new generation of sunway supercomputer. The final achievement demonstrates the effectiveness of CNN-based representation of quantum-state and brings the state-of-the-art record up to a brand-new level from both aspects of remarkable accuracy and unprecedented scales.
Mingfan Li, Junshi Chen 0003, Qingcai Jiang, Xuncheng Zhao, Rongfen Lin, Hong An, Lixin He
IEEE Trans. Parallel Distributed Syst.1
2021 swFLOW: A large-scale distributed framework for deep learning on Sunway TaihuLight supercomputer
Mingfan Li, Junshi Chen 0003, José Monsalve Diaz, Rongfen Lin, Guang R. Gao, Hong An
Inf. Sci.1
2020 Distributed deep learning system for cancerous region detection on Sunway TaihuLight
Guofeng Lv, Mingfan Li, Hong An, Junshi Chen 0003, Wenting Han, Rongfen Lin
CCF Trans. High Perform. Comput.2
2019 Degree-of-Node Task Scheduling of Fine-Grained Parallel Programs on Heterogeneous Systems
Mingfan Li, Chengfan Jia, Hong An
J. Comput. Sci. Technol.2