VLDB 2026 Research / reviewers in the wild / expert
Rongfen Lin
dblp:91/11103
· DBLP profile ↗
11ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-0321-5405ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TensorMD: Molecular Dynamics Simulation with Ab Initio Accuracy of 50 Billion AtomsabstractMolecular dynamics simulation emerges as an important area that HPC+AI helps to investigate the physical properties, with machine-learning interatomic potentials (MLIPs) being used. General-purpose machine-learning (ML) tools have been leveraged in MLIPs, but they are not perfectly matched with each other, since many optimization opportunities in MLIPs have been missed by ML tools. This inefficiency arises from the fact that HPC+AI applications work with far more computational complexity compared with pure AI scenarios. This paper has developed an MLIP, named TensorMD, independently from any ML tool. TensorMD has been evaluated on two supercomputers and scaled to 51.8 billion atoms, i.e., ~ 3× compared with state-of-the-art. Yucheng Ouyang, Ying Liu 0055, Honghui Shang, Zhenchuan Chen, Jiahao Shan, Huimin Cui, Xiaobing Feng 0002, Xingyu Gao 0003, Haifeng Song 0003, Xin Chen 0023, Rongfen Lin |
PPoPP | 13 |
| 2025 | TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic PotentialabstractAI has been integrated into HPC across various scientific fields, significantly enhancing performance. In molecular dynamics simulations, HPC+AI facilitates the investigation of atomic-scale physical properties using machine-learning interatomic potentials (MLIPs). However, general-purpose ML tools (e.g., TensorFlow) used in MLIPs are not optimally matched, leading to missed optimization opportunities due to the higher computational complexity and greater diversity of HPC+AI applications compared to pure AI scenarios. To address this, we introduce TensorMD, an MLIP independent of existing ML tools, enabling flexible optimizations that standard ML frameworks cannot support. TensorMD outperforms a state-of-the-art MLIP—winner of the 2020 Gordon Bell Prize and built on an ML tool—by 1.88 × on NVIDIA A100 GPU. Additionally, TensorMD was evaluated on two supercomputers with different architectures, achieving significantly reduced time-to-solution and supporting molecular dynamics simulations at scales beyond 50 billion atoms. Yucheng Ouyang, Ying Liu 0055, Xin Chen 0023, Honghui Shang, Zhenchuan Chen, Rongfen Lin, Xingyu Gao 0003, Jiahao Shan, Haifeng Song 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
SC | 7 |
| 2023 | Rapid simulations of atmospheric data assimilation of hourly-scale phenomena with modern neural networksabstractAtmospheric data assimilation is essential for numerical weather prediction. Ensemble data assimilation connects multiple instances of an atmospheric model through a Kalman filter-based algorithm, which is regarded as a challenging computing task today. In this work, we build a fast, low-cost, and scalable atmospheric data assimilation prototype, DIDA, for the new-generation Sunway supercomputer, including: (1) a framework that enables flexible deployment of components, and manages and optimizes data communication among modules, achieving maximum resource efficiency; (2) an accurate, robust, UNet-based surrogate model for atmospheric dynamic simulation to generate the background ensemble; (3) a batch-LETKF algorithm with high-performance eigenvalue decomposition, which is up to 7.37 times faster than existing numerical libraries while exhibiting almost linear scalability. Experimental evaluations show that our AI-integrated ensemble data assimilation prototype can complete hour-cycle assimilation in minutes, maintain linear scalability, and save an order of magnitude of computing resources, compared with the traditional method. Yiyuan Li, Xiting Ju, Qilong Jia, Yongxiao Zhou, Simeng Qian, Rongfen Lin, Bin Yang 0043, Shupeng Shi, Xin Liu 0081, Jian Tan 0005, Zhengding Hu, Limin Yan, Wei Xue 0003 |
SC | 7 |
| 2023 | 5 ExaFlop/s HPL-MxP Benchmark with Linear Scalability on the 40-Million-Core Sunway SupercomputerabstractHPL-MxP is an emerging high performance benchmark used to measure the mixed-precision computing capability of leading supercomputers. In this work, we present our efforts on the new Sunway that linearly scales the benchmark to over 40 million cores, sustains an overall mixed-precision performance exceeding 5 ExaFlop/s, and achieves over 85% of peak performance, which is the highest efficiency reached among all heterogeneous systems on the HPL-MxP list. The optimizations of our HPL-MxP implementation include the following: (1) a Two-Direction Look-Ahead and Overlap algorithm that enables overlaps of all communications with computation; (2) a multi-level process-mapping and communication scheduling method that uses the entire network as best as possible while maintaining conflict-free algorithm-flow; and (3) a CG-Fusion computing framework that eliminates up to 60% of inter-chip communications and removes the memory access bottleneck while serving both computation and communication simultaneously. This work could also provide useful insights for tuning cutting-edge applications on Sunway supercomputers as well as other heterogeneous supercomputers. Rongfen Lin, Xinhui Yuan, Wei Xue 0003, Wanwang Yin, Jienan Yao, Junda Shi, Chaobo Song, Fei Wang 0096 |
SC | 1 |
| 2022 | Large-Scale Simulation of Quantum Computational Chemistry on a New Sunway SupercomputerabstractQuantum computational chemistry (QCC) is the use of quantum computers to solve problems in computational quantum chemistry. We develop a high performance variational quantum eigensolver (VQE) simulator for simulating quantum computational chemistry problems on a new Sunway supercomputer. The major innovations include: (1) a Matrix Product State (MPS) based VQE simulator to reduce the amount of memory needed and increase the simulation efficiency; (2) a combination of the Density Matrix Embedding Theory with the MPS-based VQE simulator to further extend the simulation range; (3) A three-level parallelization scheme to scale up to 20 million cores; (4) Usage of the Julia script language as the main programming language, which both makes the programming easier and enables cutting edge performance as native C or Fortran; (5) Study of real chemistry systems based on the VQE simulator, achieving nearly linearly strong and weak scaling. Our simulation demonstrates the power of VQE for large quantum chemistry systems, thus paves the way for large-scale VQE experiments on near-term quantum computers. Honghui Shang, Li Shen 0001, Zhiqian Xu 0005, Chu Guo, Jie Liu 0069, Rongfen Lin, Yuling Yang, Zhuoya Wang, Yunquan Zhang |
SC | 9 |
| 2022 | Bridging the Gap between Deep Learning and Frustrated Quantum Spin System for Extreme-Scale Simulations on New Generation of Sunway SupercomputerabstractEfficient numerical methods are promising tools for delivering unique insights into the fascinating properties of physics, such as the highly frustrated quantum many-body systems. However, the computational complexity of obtaining the wave functions for accurately describing the quantum states increases exponentially with respect to particle number. Here we present a novel convolutional neural network (CNN) for simulating the two-dimensional highly frustrated spin-$1/2$$J_1-J_2$Heisenberg model, meanwhile the simulation is performed at an extreme scale system with low cost and high scalability. By ingenious employment of transfer learning and CNN’s translational invariance, we successfully investigate the quantum system with the lattice size up to$24\times 24$, within 30 million cores of the new generation of sunway supercomputer. The final achievement demonstrates the effectiveness of CNN-based representation of quantum-state and brings the state-of-the-art record up to a brand-new level from both aspects of remarkable accuracy and unprecedented scales. Mingfan Li, Junshi Chen 0003, Qingcai Jiang, Xuncheng Zhao, Rongfen Lin, Hong An, Lixin He |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2021 | TensorKMC: kinetic Monte Carlo simulation of 50 trillion atoms driven by deep learning on a new generation of Sunway supercomputerabstractThe atomic kinetic Monte Carlo method plays an important role in multi-scale physical simulations because it bridges the micro and macro worlds. However, its accuracy is limited by empirical potentials. We therefore propose herein a triple-encoding algorithm and vacancy-cache mechanism to efficiently integrate ab initio neural network potentials (NNPs) with AKMC and implement them in our TensorKMC codes. We port our program to SW26010-pro and innovate a fast feature operator and a big fusion operator for the NNPs for fully utilizing the powerful heterogeneous computing units of the new-generation Sunway supercomputer. We further optimize memory usage. With these improvements, TensorKMC can simulate up to 54 trillions of atoms and achieve excellent strong and weak scaling performance up to 27,456,000 cores. Honghui Shang, Xin Chen 0023, Xingyu Gao 0003, Rongfen Lin, Lei Xu 0023, Leilei Zhu, Fei Wang 0096, Yunquan Zhang, Haifeng Song 0003 |
SC | 4 |
| 2021 | Extreme-scale ab initio quantum raman spectra simulations on the leadership HPC system in ChinaabstractRaman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations, are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, full QM simulations have an extremely high computational cost and remain challenging. In this work, robust new algorithms and advanced implementations on many-core architectures are employed to enable fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of realistic biological systems containing up to 3006 atoms, with excellent strong and weak scaling. Up to a performance of 468.5 PFLOP/s in double-precision and 813.7 PLOPS/s in mixed-half precision is achieved on the new-generation Sunway high-performance computing system, suggesting the potential for new applications of the QM approach to biological systems. Honghui Shang, Yunquan Zhang, You Fu, Yingxiang Gao, Yangjun Wu, Xiaohui Duan, Rongfen Lin, Xin Liu 0081, Ying Liu 0055, Dexun Chen |
SC | 9 |
| 2021 | swFLOW: A large-scale distributed framework for deep learning on Sunway TaihuLight supercomputer
Mingfan Li, Junshi Chen 0003, José Monsalve Diaz, Rongfen Lin, Guang R. Gao, Hong An |
Inf. Sci. | 6 |
| 2020 | Distributed deep learning system for cancerous region detection on Sunway TaihuLight
Guofeng Lv, Mingfan Li, Hong An, Junshi Chen 0003, Wenting Han, Rongfen Lin |
CCF Trans. High Perform. Comput. | 9 |
| 2017 | Towards Highly Efficient DGEMM on the Emerging SW26010 Many-Core ProcessorabstractThe matrix-matrix multiplication is an essential building block that can be found in various scientific and engineering applications. High-performance implementations of the matrix-matrix multiplication on state-of-the-art processors may be of great importance for both the vendors and the users. In this paper, we present a detailed methodology of implementing and optimizing the double-precision general format matrix-matrix multiplication (DGEMM) kernel on the emerging SW26010 processor, which is used to build the Sunway TaihuLight supercomputer. We propose a three level blocking algorithm to orchestrate data on the memory hierarchy and expose parallelism on different hardware levels, and design a collective data sharing scheme by using the register communication mechanism to exchange data efficiently among different cores. On top of those, further optimizations are done based on a data-thread mapping method for efficient data distribution, a double buffering scheme for asynchronous DMA data transfer, and an instruction scheduling method for maximizing the pipeline usage. Experiment results show that the proposed DGEMM implementation can fully exploit the unique hardware features provided by SW26010 and can sustain up to 95% of the peak performance. Lijuan Jiang, Chao Yang 0002, Yulong Ao, Wanwang Yin, Wenjing Ma, Qiao Sun 0005, Fangfang Liu 0004, Rongfen Lin |
ICPP | 8 |