EDBT 2026 Demo / reviewers in the wild / expert
Xin Chen 0023
dblp:24/1518-23
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-0562-0319ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TensorMD: Molecular Dynamics Simulation with Ab Initio Accuracy of 50 Billion AtomsabstractMolecular dynamics simulation emerges as an important area that HPC+AI helps to investigate the physical properties, with machine-learning interatomic potentials (MLIPs) being used. General-purpose machine-learning (ML) tools have been leveraged in MLIPs, but they are not perfectly matched with each other, since many optimization opportunities in MLIPs have been missed by ML tools. This inefficiency arises from the fact that HPC+AI applications work with far more computational complexity compared with pure AI scenarios. This paper has developed an MLIP, named TensorMD, independently from any ML tool. TensorMD has been evaluated on two supercomputers and scaled to 51.8 billion atoms, i.e., ~ 3× compared with state-of-the-art. Yucheng Ouyang, Ying Liu 0055, Honghui Shang, Zhenchuan Chen, Jiahao Shan, Huimin Cui, Xiaobing Feng 0002, Xingyu Gao 0003, Haifeng Song 0003, Xin Chen 0023, Rongfen Lin |
PPoPP | 12 |
| 2025 | TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic PotentialabstractAI has been integrated into HPC across various scientific fields, significantly enhancing performance. In molecular dynamics simulations, HPC+AI facilitates the investigation of atomic-scale physical properties using machine-learning interatomic potentials (MLIPs). However, general-purpose ML tools (e.g., TensorFlow) used in MLIPs are not optimally matched, leading to missed optimization opportunities due to the higher computational complexity and greater diversity of HPC+AI applications compared to pure AI scenarios. To address this, we introduce TensorMD, an MLIP independent of existing ML tools, enabling flexible optimizations that standard ML frameworks cannot support. TensorMD outperforms a state-of-the-art MLIP—winner of the 2020 Gordon Bell Prize and built on an ML tool—by 1.88 × on NVIDIA A100 GPU. Additionally, TensorMD was evaluated on two supercomputers with different architectures, achieving significantly reduced time-to-solution and supporting molecular dynamics simulations at scales beyond 50 billion atoms. Yucheng Ouyang, Ying Liu 0055, Xin Chen 0023, Honghui Shang, Zhenchuan Chen, Rongfen Lin, Xingyu Gao 0003, Jiahao Shan, Haifeng Song 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
SC | 4 |
| 2024 | Multilevel Load Balancing Algorithm for Domestic Heterogeneous Manycore ArchitectureabstractLoad imbalance often occurs in particle-in-cell simulations on parallel computing, which seriously affects the efficiency of applications. Due to the characteristics of multilevel parallelism and communication asymmetry of compute nodes in domestic heterogeneous manycore architecture, the impact of load imbalance is more prominent. The paper proposes a multilevel load-balancing algorithm for domestic heterogeneous manycore architecture. Inside the supernode, computing tasks are redivided based on manycore acceleration. Between the supernodes, a greedy-based communication mode is designed to minimize communication across supernodes. The experimental results show that the proposed algorithm achieves almost ideal dynamic load balance, and improves the performance of the evaporation module in two-phase flow simulation by 10.9-19.7 times for the 50 million-sized grid. Xin Chen 0023, Xin Liu 0081 |
ISPA | 2 |
| 2023 | Scalability and efficiency challenges for the exascale supercomputing system: practice of a parallel supporting environment on the Sunway exascale prototype systemabstractWith the continuous improvement of supercomputer performance and the integration of artificial intelligence with traditional scientific computing, the scale of applications is gradually increasing, from millions to tens of millions of computing cores, which raises great challenges to achieve high scalability and efficiency of parallel applications on super-large-scale systems. Taking the Sunway exascale prototype system as an example, in this paper we first analyze the challenges of high scalability and high efficiency for parallel applications in the exascale era. To overcome these challenges, the optimization technologies used in the parallel supporting environment software on the Sunway exascale prototype system are highlighted, including the parallel operating system, input/output (I/O) optimization technology, ultra-large-scale parallel debugging technology, 10-million-core parallel algorithm, and mixed-precision method. Parallel operating systems and I/O optimization technology mainly support large-scale system scaling, while the ultra-large-scale parallel debugging technology, 10-million-core parallel algorithm, and mixed-precision method mainly enhance the efficiency of large-scale applications. Finally, the contributions to various applications running on the Sunway exascale prototype system are introduced, verifying the effectiveness of the parallel supporting environment design. Xiaobin He, Xin Chen 0023, Xin Liu 0081, Dexun Chen, Yuling Yang, Yunlong Feng, Longde Chen, Xiaona Diao, Zuoning Chen |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2023 | Redesigning OpenKMC for Multi-Component Trillion-Atom Simulations on the New Sunway SupercomputerabstractThe atomic kinetic Monte Carlo method plays an important role in material simulations by connecting the microscale mechanism with macroscale evolution. However, the long-time simulation of multi-component materials is highly challenging because it demands significant computing resources. With the advent of exascale computing, ultra-high computing power can enable kinetic Monte Carlo (KMC) simulations. In this paper, we deeply optimize OpenKMC for the new-generation Sunway supercomputer. This includes optimizing the memory access for the SW39000 architecture, eliminating various redundant computations at growing scales, and proposing a communication strategy for heterogeneous platforms. In addition, we expanded OpenKMC's simulation for multi-component alloys. Finally, the acceleration framework can produces a$37\times$performance enhancement on the Sunway platform. Furthermore, when powered by 10 million cores, our program can perform trillion-atom simulations of complex multi-component alloys with 85% parallel efficiency. Lei Xu 0023, Honghui Shang, Xin Chen 0023, Yunquan Zhang, Xingyu Gao 0003, Haifeng Song 0003 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Increasing the Efficiency of Massively Parallel Sparse Matrix-Matrix Multiplication in First-Principles Calculation on the New-Generation Sunway SupercomputerabstractThe first-principles approach based on density-functional theory (DFT)/density-functional perturbation theory (DFPT) is widely used in calculations of the systems’ ground state energy, response properties (e.g., polarizability, phonon dispersions) and is playing an increasingly important role in chemistry, physics and materials science. For the large-scale calculations, the computation of the density matrix/response density matrix in DFT/DFPT has become the main performance bottleneck. One of the solutions is using the linear scaling method to get the density matrix and response density matrix. Here a massively parallel medium sparse matrix-matrix multiplication algorithm is designed for first-principle calculations and implemented on the new-generation Sunway supercomputer. Experiments show that the proposed method has obvious performance advantages compared to the original parallel version under moderate sparsity. The computing cores scale to 3,900,000 with strong scalability of 77.3$\%$. Xin Chen 0023, Yingxiang Gao, Honghui Shang, Zhiqian Xu 0005, Xin Liu 0081, Dexun Chen |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | TensorKMC: kinetic Monte Carlo simulation of 50 trillion atoms driven by deep learning on a new generation of Sunway supercomputerabstractThe atomic kinetic Monte Carlo method plays an important role in multi-scale physical simulations because it bridges the micro and macro worlds. However, its accuracy is limited by empirical potentials. We therefore propose herein a triple-encoding algorithm and vacancy-cache mechanism to efficiently integrate ab initio neural network potentials (NNPs) with AKMC and implement them in our TensorKMC codes. We port our program to SW26010-pro and innovate a fast feature operator and a big fusion operator for the NNPs for fully utilizing the powerful heterogeneous computing units of the new-generation Sunway supercomputer. We further optimize memory usage. With these improvements, TensorKMC can simulate up to 54 trillions of atoms and achieve excellent strong and weak scaling performance up to 27,456,000 cores. Honghui Shang, Xin Chen 0023, Xingyu Gao 0003, Rongfen Lin, Lei Xu 0023, Leilei Zhu, Fei Wang 0096, Yunquan Zhang, Haifeng Song 0003 |
SC | 2 |
| 2021 | Accelerating all-electron ab initio simulation of raman spectra for biological systemsabstractRaman spectroscopy provides chemical and compositional information that can serve as a structural fingerprint for various materials. Therefore, simulations of Raman spectra, including both quantum perturbation analyses and ground-state calculations are of significant interest. However, highly accurate full quantum mechanical (QM) simulations of Raman spectra have previously been confined to small systems. For large systems such as biological materials, the computational cost of full QM simulations is extremely high, and their extension to such systems remains challenging. In the work described here, by employing robust new algorithms and advances in implementation for the many-core architectures, we are able to perform fast, accurate, and massively parallel full ab initio simulations of the Raman spectra of biological systems with excellent strong and weak scaling, thereby providing a starting point for applying QM approaches to structural studies of such systems. Honghui Shang, Yunquan Zhang, Ying Liu 0055, Mingchuan Wu, Yangjun Wu, Di Wei, Huimin Cui, Xin Liu 0081, Fei Wang 0096, Yuxi Ye, Yingxiang Gao, Shuang Ni, Xin Chen 0023, Dexun Chen |
SC | 15 |
| 2018 | Structural total least squares algorithm for locating multiple disjoint sources based on AOA/TOA/FOA in the presence of system errorabstractSingle-station passive localization technology avoids the complex time synchronization and information exchange between multiple observatories, and is increasingly important in electronic warfare. Based on a single moving station localization system, a new method with high localization precision and numerical stability is proposed when the measurements from multiple disjoint sources are subject to the same station position and velocity displacement. According to the available measurements including the angle-of-arrival (AOA), time-of-arrival (TOA), and frequency-of-arrival (FOA), the corresponding pseudo linear equations are deduced. Based on this, a structural total least squares (STLS) optimization model is developed and the inverse iteration algorithm is used to obtain the stationary target location. The localization performance of the STLS localization algorithm is derived, and it is strictly proved that the theoretical performance of the STLS method is consistent with that of the constrained total least squares method under first-order error analysis, both of which can achieve the Cramér-Rao lower bound accuracy. Simulation results show the validity of the theoretical derivation and superiority of the new algorithm. Xin Chen 0023, Ding Wang 0003, Rui-rui Liu 0001, Jiexin Yin, Ying Wu 0002 |
Frontiers Inf. Technol. Electron. Eng. | 1 |