EDBT 2026 Demo / reviewers in the wild / expert
Xinming Qin
dblp:260/9001
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Million-Atom Ab Initio Electron Dynamics: Discontinuous Galerkin Real-Time Time-Dependent Density Functional TheoryabstractOver the past decades, first-principles real-time time dependent density functional theory(rt-TDDFT) simulations have been limited to systems with only thousands of atoms. We propose a novel method based on the discontinuous Galerkin adaptive local basis, significantly reducing global communication in rt-TDDFT. We further introduce a tensor compression technique that leverages basis locality to avoid repeated evaluation of multi-center integrals in hybrid functionals, greatly reducing computational cost. To overcome the projection bottleneck in our basis sets, we design a fused Gemm-Reduce operation that achieves several times higher floating-point efficiency than standard BLAS combination. Our implementation reaches 34.8% of theoretical peak performance on 524,288 CGs of the New Sunway supercomputer and simulates electronic dynamics of systems with over one million atoms for both local-semi-local and hybrid functionals. This work improves computational scale by two orders of magnitude, opening new possibilities for exploring ultrafast dynamics in large-scale materials and nanophotonic devices. Junwei Feng, Junshi Chen 0003, Xinming Qin, Lingyun Wan, Wentiao Wu, Bingkun Hou, Yexuan Lin, Zechuan Zhang, Weile Jia, Hong An, Jinlong Yang 0003, Wei Hu 0006 |
SC | 5 |
| 2025 | Considering Student's Cognitive Abilities-Entity Level and Entirety Level Sentiment Analysis of Teaching EvaluationabstractStudent Teaching Feedback Sentiment Analysis (STFSA) plays a crucial role in evaluating teaching effectiveness, and its analysis results are influenced by both teaching process organization and students’ course cognition. However, current sentiment analysis of students’ feedback faces the challenges of integrating fine-grained entity-level sentiment and multi-polar sentiment in overall sentiment analysis. Therefore, we propose a novel sentiment analysis strategy for Entity-level and Entirety-level Teaching evaluation Sentiment analysis that considers students’ Cognitive abilities (EETSC). This strategy utilizes two-level networks: one for the entity level network based on dual attention mechanism to conduct fine-grained sentiment analysis of students’ feedback, and the other for the overall sentiment analysis network that integrates students’ personalized cognitive abilities to adjust the sentiment characteristics of different entities and obtain more accurate and reasonable emotions at the overall level. Experimental results show that, compared to the state-of-the-art baseline, EETSC achieves an accuracy of 86.34% and an F1 score of 77.13% in the recognition of six types of teaching entity sentiments, representing improvements of 2.19% and 2.18%, respectively. For entirety-level sentiment recognition, EETSC achieves an accuracy of 87.36% and an F1 score of 78.02%, with improvements of 0.76% and 1.21%. Further experimental analysis indicates that EETSC can alleviate the problem of sentiment polarity conflicts in teaching evaluations and provides a solution for integrating students’ cognitive states into teaching sentiment analysis in the field of educational natural language processing. Xinming Qin |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2025 | PWDFT-SW: Extending the Limit of Plane-Wave DFT Calculations to 16K Atoms on the New Sunway SupercomputerabstractFirst-principles density functional theory (DFT) with plane wave (PW) basis set is the most widely used method in quantum mechanical material simulations due to its advantages in accuracy and universality. However, a perceived drawback of PW-based DFT calculations is their substantial computational cost and memory usage, which currently limits their ability to simulate large-scale complex systems containing thousands of atoms. This situation is exacerbated in the new Sunway supercomputer, where each process is limited to a mere 16 GB of memory. Herein, we present a novel parallel implementation of plane wave density functional theory on the new Sunway supercomputer (PWDFT-SW). PWDFT-SW fully extracts the benefits of Sunway supercomputer by extensively refactoring and calibrating our algorithms to align with the system characteristics of the Sunway system. Through extensive numerical experiments, we demonstrate that our methods can substantially decrease both computational costs and memory usage. Our optimizations translate to a speedup of 64.8x for a physical system containing 4,096 silicon atoms, enabling us to push the limit of PW-based DFT calculations to large-scale systems containing 16,384 carbon atoms. Qingcai Jiang, Zhenwei Cao, Junshi Chen 0003, Xinming Qin, Wei Hu 0006, Hong An, Jinlong Yang 0003 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | DB-SpGEMM: A Massively Distributed Block-Sparse Matrix-Matrix Multiplication for Linear-Scaling DFT CalculationsabstractLinear-scaling <?TeX $\mathcal {O}(N)$?> Math 1 density functional theory (DFT) represents a significant advancement in the field of computational materials science, especially for simulations of large systems where traditional cubic-scaling methods become computationally prohibitive. The core operation in <?TeX $\mathcal {O}(N)$?> Math 2 methods is sparse general matrix-matrix multiplication (SpGEMM), which is the major performance bottleneck. To enhance the computational efficiency of SpGEMM, it is crucial to consider the inherent sparse pattern of these matrices. Targeting block-sparse matrices with moderate block sizes and regular block shapes, we have developed a distributed block-sparse matrix-matrix multiplication (DB-SpGEMM) algorithm for large-scale DFT calculations. Through deep optimizations in distributed matrix storage, computational task decomposition, asynchronous task scheduling, and load balancing, we have implemented a linear-scaling method based on this algorithm within the discontinuous Galerkin density functional theory (DGDFT). On the new Sunway supercomputer, our approach achieves a 8 ∼ 10x speedup compared to the original version on monolayer phosphorene systems, and demonstrates superior scalability. Junshi Chen 0003, Yang Zhao 0040, Longsheng Song, Xinming Qin, Hong An |
ICPP | 5 |
| 2024 | Enabling 13K-Atom Excited-State GW Calculations via Low-Rank Approximations and HPC on the New Sunway SupercomputerabstractGW approximation is a powerful approach to accurately describe the excited-state of semiconductors. However, GW incurs high computational cost $\mathcal{O}\left(N^{4}\right)$ and large memory usage $\mathcal{O}\left(N^{3}\right)$, limiting its applications to thousands of (2,742) atoms even on leadership supercomputers. Herein we present a massively parallel implementation of accurate and efficient cubic-scaling plane-wave GW calculations by using low-rank approximations and high-performance computing on leadership supercomputers. By using a series of low rank approximations, we can reduce the expensive GW calculations to the cubic-scaling computational cost $\mathcal{O}\left(N^{3}\right)$ and quadratic memory usage $\mathcal{O}\left(N^{2}\right)$. With the help of parallel and communication optimization, the plane-wave GW calculations gain an overall speedup of over 70x and efficiently scale up to 13,824 atoms within a few minutes using 449,280 cores on new Sunway supercomputer. This accomplishment paves the way for excited-state quantum mechanical material simulations at mesoscopic scale (10K atoms) and for the design of next-generation semiconductor devices. Wentiao Wu, Zhengbang Zhou, Qingcai Jiang, Junwei Feng, Xinming Qin, Huanhuan Ma, Zhenwei Cao, Junshi Chen 0003, Xinyong Meng, Bingkun Hou, Yuanfan Xiong, Linhao Wang, Yixuan Sun, Hong An, Jinlong Yang 0003, Wei Hu 0006 |
SC | 5 |
| 2024 | Extending the limit of LR-TDDFT on two different approaches: Numerical algorithms and new Sunway heterogeneous supercomputerabstractFirst-principles time-dependent density functional theory (TDDFT) is a powerful tool to accurately describe the excited-state properties of molecules and solids in condensed matter physics , computational chemistry, and materials science. However, a perceived drawback in TDDFT calculations is its ultrahigh computational cost O ( N 5 ∼ N 6 ) and large memory usage O ( N 4 ) especially for plane-wave basis set, confining its applications to large systems containing thousands of atoms. Here, we present a massively parallel implementation of linear-response TDDFT (LR-TDDFT) and accelerate LR-TDDFT in two different aspects: (1) numerical algorithms on the X86 supercomputer and (2) optimizations on the heterogeneous architecture of the new Sunway supercomputer. Furthermore, we carefully design the parallel data and task distribution schemes to accommodate the physical nature of different computation steps. By utilizing these two different methods, our implementation can gain an overall speedup of 10x and 80x and efficiently scales to large systems up to 4096 and 2744 atoms within dozens of seconds. Qingcai Jiang, Zhenwei Cao, Xinhui Cui, Lingyun Wan, Xinming Qin, Huanqi Cao, Hong An, Junshi Chen 0003, Jie Liu 0069, Wei Hu 0006, Jinlong Yang 0003 |
Parallel Comput. | 5 |
| 2023 | High performance computing for first-principles Kohn-Sham density functional theory towards exascale supercomputers
Xinming Qin, Junshi Chen 0003, Zhaolong Luo, Lingyun Wan, Jielan Li, Shizhe Jiao, Qingcai Jiang, Wei Hu 0006, Hong An, Jinlong Yang 0003 |
CCF Trans. High Perform. Comput. | 1 |
| 2022 | Accelerating Parallel First-Principles Excited-State Calculation by Low-Rank Approximation with K-Means ClusteringabstractFirst-principles time-dependent density functional theory (TDDFT) is a powerful tool to accurately describe the excited-state properties of molecules and solids in condensed matter physics, computational chemistry and materials science. However, a perceived drawback in TDDFT calculations is its ultrahigh computational cost and large memory usage especially for plane-wave basis set, confining its applications to large systems containing thousands of atoms. Here, we present a massively parallel implementation of linear-response TDDFT (LR-TDDFT) and reduce the complexity to by combining K-Means clustering based low-rank approximation with iterative eigensolve algorithm. Furthermore, we carefully design the parallel data and task distribution schemes to accommodate with the physical nature in different steps of the computation, also, several optimization methods are employed to effectively handle the matrix operations and data communications of constructing and diagonalizing the LR-TDDFT Hamiltonian. In particular, our method can significantly reduce the cost of computation and memory by nearly 2 orders of magnitude compared to conventional LR-TDDFT calculations. Numerical results demonstrate that our implementation can gain an overall speedup of 10x and efficiently scale up to 12,288 CPU cores for large systems up to 4,096 atoms within dozens of seconds. Qingcai Jiang, Jielan Li, Junshi Chen 0003, Xinming Qin, Lingyun Wan, Jinlong Yang 0003, Jie Liu 0069, Wei Hu 0006, Hong An |
ICPP | 4 |
| 2022 | 2.5 Million-Atom Ab Initio Electronic-Structure Simulation of Complex Metallic Heterostructures with DGDFTabstractOver the past three decades, ab initio electronic structure calculations of large, complex and metallic systems are limited to tens of thousands of atoms in computational accuracy and efficiency on leadership supercomputers. We present a massively parallel discontinuous Galerkin density functional theory (DGDFT) implementation, which adopts adaptive local basis functions to discretize the Kohn-Sham equation, resulting in a block-sparse Hamiltonian matrix. A highly efficient pole expansion and selected inversion (PEXSI) sparse direct solver is implemented in DGDFT to achieve O(N1.5) scaling for quasi two-dimensional systems. DGDFT allows us to compute the electronic structures of complex metallic heterostructures with 2.5 million atoms (17.2 million electrons) using 35.9 million cores on the new Sunway supercomputer. The peak performance of PEXSI can achieve 64 PFLOPS (~5% of theoretical peak), which is un-precedented for sparse direct solvers. This accomplishment paves the way for quantum mechanical simulations into mesoscopic scale for designing next-generation electronic devices. Wei Hu 0006, Hong An, Zhuoqiang Guo, Qingcai Jiang, Xinming Qin, Junshi Chen 0003, Weile Jia, Chao Yang 0001, Zhaolong Luo, Jielan Li, Wentiao Wu, Guangming Tan, Dongning Jia, Qinglin Lu, Yeqi Huang, Liyi Wang, Jinlong Yang 0003 |
SC | 5 |