Qi Du

dblp:158/3378 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Exploring the performance of CP2K simulations on the CPU-GPDSP Fusion intra-heterogeneous HPC system
Qi Du
Future Gener. Comput. Syst.1
2025 Scaling Deep Learning Molecular Dynamics to 500M Atoms on 4096-Node ARMv8 Clusters
abstract
Molecular dynamics (MD) simulations are essential tools for investigating large-scale molecular systems, yet achieving high performance and scalability on CPU-based architectures remains challenging. In this study, we present a highly optimized framework based on DeepMD-kit for conducting 500 millionatom MD simulations on an ARMv8 SVE high-performance computing (HPC) system. Key optimizations include leveraging OpenMP for multi-threaded acceleration of DeepMD-kit and utilizing the ARMv8 SVE instruction set to optimize doubleprecision matrix multiplication in PyTorch. These enhancements enable single ARMv8 SVE 64-core processors to achieve 1.3x the training performance of NVIDIA V100 GPU, and two ARMv8 SVE 64-core processors to achieve 1.05x the inference performance of NVIDIA V100 GPU. Leveraging this optimized framework, we achieve large-scale MD simulations across 4,096 computing nodes.
Qi Du, Feng Wang 0050, Chengkun Wu, Han Wang 0006, Yongpeng Liu, Zhaoyin Zhou, Kenli Li 0001
CLUSTER1
2025 Improving LAMMPS performance for molecular dynamic simulation on large-scale HPC systems
abstract
Abstract Large-scale atomic/molecular massively parallel simulator (LAMMPS) is a prevalent software package employed for molecular dynamics simulations, enabling the study of materials at the atomic and molecular scale. Its performance is paramount in numerous industrial applications, driving the need for ongoing enhancements in simulation speed and parallel efficiency. Previous works heavily rely on hardware accelerators, which lead to limited parallel and high costs. To address this, this work optimizes the message passing interface (MPI) and memory copy functions, while deploying LAMMPS on high-performance computing (HPC) systems. We propose a new adaptive broadcast algorithm to improve the parallelism efficiency of the interconnect topology. We also discuss how to realize the mutual hiding of computation and communication of the Packing algorithm in LAMMPS, and optimize the memory copy function and MPI operators to facilitate the execution of the program. The resulting components are integrated into the MPICH4 software and deployed on the MT-3000 HPC system. The experimental results show a significant performance improvement, with up to four orders of magnitude speedup on 1024, and more than 90% parallel efficiencies, demonstrating the effectiveness of our proposed optimization scheme. The adaptive broadcast algorithm and the portability of computation and communication hiding are also discussed. The adaptive broadcast algorithm is applied to SPEC MPI2007, and the average performance improvement is 23.91 and 27.29% on ARMv8 cluster and x86_64 cluster, respectively.
Qi Du, Jinlin Chen
Comput. J.1
2025 Parallelization Strategies for DeepMD-Kit Using OpenMP: Enhancing Efficiency in Machine Learning-Based Molecular Simulations
abstract
DeepMD-kit enables deep learning-based molecular dynamics (MD) simulations that require efficient parallelization to leverage modern HPC architectures. In this work, we optimize DeepMD-kit using advanced OpenMP strategies to improve scalability and computational efficiency on an ARMv8 processor-based server. Our optimizations include data parallelism for neural network inference, force calculation acceleration, NUMAaware memory management, and synchronization reductions, leading to up to 4.1× speedup and 82% higher memory band-width efficiency compared to the baseline implementation. Strong scaling analysis demonstrates superlinear speedup at mid-range core counts, with improved workload balancing and vectorized computations. However, challenges remain at ultra-large scales due to increasing synchronization overhead.
Qi Du, Feng Wang 0050, Chengkun Wu
IEEE Trans. Computers1
2024 Exploring Natural Language Processing Model Acceleration in Molecular Dynamics Simulation Using High-Performance Computing and Machine Learning
abstract
In molecular dynamics simulations, the integral methods used to calculate molecular trajectories, such as the Verlet integral method, involve significant computational costs. On the premise of ensuring the accuracy of simulation, how to effectively reduce the amount of computation is always a challenging problem. This paper applies Natural Language Processing models to simulate molecular motion trajectories in molecular dynamics, integrating the MPI programming model and machine learning on a high-performance computing platform to achieve MPI parallelization during model training and inference. The results indicate that different types of neural network architectures have a significant impact on inference performance and accuracy. When using the deep spatio-temporal networks model, the mean absolute error is 0.0032. At an atomic scale of 3M, the parallel efficiency with 1024 computational nodes exceeds 90%, demonstrating excellent parallel performance.
Qi Du, Feng Wang 0050, Heng Wan, Chengkun Wu
BIBM1
2022 MPI parameter optimization during debugging phase of HPC system
Qi Du
J. Supercomput.1
2021 Placement for Wafer-Scale Deep Learning Accelerator
abstract
To meet the growing demand from deep learning applications for computing resources, accelerators by ASIC are necessary. A wafer-scale engine (WSE) is recently proposed [1], which is able to simultaneously accelerate multiple layers from a neural network (NN). However, without a high-quality placement that properly maps NN layers onto the WSE, the acceleration efficiency cannot be achieved. Here, the WSE placement resembles the traditional ASIC floor plan problem of placing blocks onto a chip region, but they are fundamentally different. Since the slowest layer determines the compute time of the whole NN on WSE, a layer with a heavier workload needs more computing resources. Besides, locations of layers and protocol adapter cost of internal 10 connections will influence inter-layer communication overhead. In this paper, we propose GigaPlacer to handle this new challenge. A binary-search-based framework is developed to obtain a minimum compute time of the NN. Two dynamic-programming-based algorithms with different optimizing strategies are integrated to produce legal placement. The distance and adapter cost between connected layers will be further minimized by some refinements. Compared with the first place of the ISPD2020 Contest, GigaPlacer reduces the contest metric by up to 6.89% and on average 2.09%, while runs 7.23X faster.
Benzheng Li, Qi Du, Dingcheng Liu, Jingchong Zhang, Gengjie Chen, Hailong You
ASP-DAC2
2020 Improving Analysis of Automatic Distribution Changes for Power Grid
abstract
In view of the actual demand of power supply enterprises for the construction of distribution network, it is necessary to integrate the two kinds of business, namely the abnormal operation of distribution network equipment and the operation and monitoring of distribution network, through the integration of power network sensor network and other technologies, the operation and distribution data and the integration of “big data” technology to carry out integrated management of the entire distribution network dispatching and production business. Based on PMS2.5 and GIS2.0, we propose a new scheme to analyse distribution changes for power grid, with the help of the parameter comparison analysis, the mutual authentication between the two systems construction of distribution network equipment move detection function. It can realize electricity information area over, demolition, shut down the usage scenarios and system parameter changes between the data management, at the same time auxiliary by automatic matching method, improved the camp with penetration data accuracy.
Meile Shi, Qi Du
MSN4
2018 A unified scheme of text localization and structured data extraction for joint OCR and data mining
abstract
Both text detection and structured data extraction are imperative in an optical character recognition (OCR) processing pipeline. Text detection, especially for indistinct, diverse, multi-language text regions, is one of the most challenging tasks in computer vision and has attracted increasing attention recently. Moreover, although there are some studies in data mining related to structured data extraction, it has not received its deserved attention as one of important steps in OCR. The previous methods for structural data extraction, including layout template-based, rule-based, and natural language processing (NLP)-based methods, usually leads to either inaccurate results or complex modules. In this paper, we integrate text detection and structured data extraction into a unified deep learning-based Image Text Extraction (ITE) scheme. Our ITE is an end-to-end trainable model and able to handle multi-scale and multi-lingual text in a single process. Experiments on large-scale real-world passport and medical receipt datasets have demonstrated the superiority of the proposed method in terms of both effectiveness and efficiency.
Yibin Ye, Shenggao Zhu, Jing Wang 0221, Qi Du, Yezhang Yang, Dandan Tu, Lanjun Wang, Jiebo Luo 0001
IEEE BigData4