VLDB 2026 Research / reviewers in the wild / expert
Shaokang Du
dblp:372/9829
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0004-0540-0959ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MatrixFold: Unleashing Manycore CPUs with Outer-Product Units for Mixed-Precision AlphaFold Inference
Shaokang Du, Hailong Yang 0002, Xin You 0001, Baojian Zhou, Depei Qian 0001 |
Euro-Par (2) | 1 |
| 2025 | Accelerating the Cryo-EM Structure Determination in RELION on Modern Many-Core CPUabstractRELION is a widely-used software suite for cryoelectron microscopy (cryo-EM) single-particle analysis (SPA), yet its performance optimization has primarily focused on x86 CPUs and NVIDIA GPUs. In this work, we present the first systematic effort to optimize RELION on modern many-core CPUs. Through detailed performance analysis, we identify critical bottlenecks across RELION's major computational stages. We then apply a set of software- and hardware-aware optimizations, including vectorization optimization, process and thread configurations tuning, algorithm optimization, lock optimization, memory affinity optimization, and computation redundancy optimization. Our optimized version achieves significant speedups and exhibits better scalability than the original RELION across all stages. Notably, it outperforms a single NVIDIA A100 GPU on the complete SPA workflow, achieving a$2.22 \times$speedup on the SPA dataset and a$1.13 \times$speedup on the RELION Benchmark dataset. Validation experiments further confirm that our optimizations preserve the reconstruction accuracy, demonstrating the potential of specific CPU architectures as a competitive and efficient platform for cryo-EM data processing. Kelun Lei, Hailong Yang 0002, Jia Yuan, Shaokang Du, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
HPCC | 4 |
| 2025 | Accelerating the Martian Atmospheric Simulation of GoMars Model with Multi-GPUsabstractMars exploration is at the forefront of space science, which demands robust computational models to decipher its atmospheric dynamics. In this work, we present a significant advancement in computational efficiency for the GoPlanetMars (GoMars), a state-of-the-art Martian atmospheric model. By leveraging the parallel processing capabilities of Graphics Processing Units (GPUs), we accelerate the dynamic core of the GoMars model on multiple NVIDIA A800 GPUs. Through comprehensive performance analysis of GoMars, we optimize both parallel computation and communication patterns to leverage the computational power of multiple GPUs fully, achieving performance comparable to that of a thousand-core CPU cluster. Our evaluation results demonstrate that the GPU-accelerated GoMars model maintains the same level of precision as the native CPU-based implementation, while achieving a substantial speedup, making it a viable solution for high-performance Martian atmospheric simulations. Guofan Yu, Haoran Kong, Xin You 0001, Hailong Yang 0002, Shaokang Du, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
HPCC | 5 |
| 2025 | OVERT: Orchestrating Vector-Scalar Execution for Efficient SpMV on Modern CPUsabstractSparse Matrix-Vector Multiplication (SpMV) is a key operation in many applications, and optimizing its performance is crucial for achieving high computational efficiency. Existing efforts have optimized SpMV performance on CPUs with corresponding sparse matrix formats adopted. However, the performance of existing SpMV implementations primarily focuses on maximizing hardware’s vector unit usage, neglecting the potential for exploiting idle scalar units simultaneously. To address such limitation, we propose OVERT, a new storage format of sparse matrix designed to exploit both vector and scalar execution units on modern CPUs for accelerating SpMV performance. OVERT, containing two format variants (OVERT-S and OVERT-E), outperforms existing formats by partitioning the matrix into multiple data panels, which can efficiently utilize vector and scalar units. Moreover, we propose an effective format selection model that dynamically chooses the optimal format variant from OVERT according to the characteristics of the input matrix. Experimental results on SuiteSparse show that OVERT achieves an average speedup of 3.91 × against Intel MKL on X86 CPU and an average speedup of 1.24 × against ArmPL on ARM CPU. Kelun Lei, Hailong Yang 0002, Kaige Zhang 0002, Shaokang Du, Marc Casas, Yufan Xu 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
ICPP | 4 |
| 2025 | Accelerating Complex Stencil Computations with Adaptive Fusion StrategyabstractStencil computation is an important computational pattern widely utilized in various scientific applications, such as image processing, climate forecasting, and fluid dynamics.With the increasing demands for higher precision by scientific applications, stencil computations have become complex, containing a set of dependent stencil operators that may process multiple input grids.These stencils are referred to as complex stencils.For complex stencils, optimizing individual stencil operators is insufficient, and there is significant interest in developing optimization approaches across stencil operators.Existing stencil optimizations or compilers adopt the producer-consumer fusion of stencil operators to Hailong Yang 0002, Shaokang Du, Yufan Xu 0001, Qingxiao Sun, Xuning Liang, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
ICS | 4 |
| 2025 | Zero-Value Code Specialization via Profile-Guided Control Data Flow AnalysisabstractZero-value propagation is a common phenomenon in modern programs, where redundant operations caused by zero-values can severely impact performance. Since zero-values are often generated dynamically at runtime, eliminating such redundancies through static analysis alone is challenging. In this paper, we propose an efficient static control data flow analysis algorithm to identify redundancies resulting from zero-value propagation. Based on this algorithm, we design and implement ZeroSpec, a fully automated profile-guided code optimizer that detects zero-values at runtime and specializes fast paths for them. To maximize performance gains, ZeroSpec also employs a fine-grained cost model that evaluates the optimization potential of individual zero-value instructions to guide the construction of targeted optimization regions. Evaluation on SPEC CPU2017, NPB and real-world applications demonstrates the effectiveness of ZeroSpec, achieving a maximum performance speedup of 1.31 ×. Shaokang Du, Kelun Lei, Xin You 0001, Hailong Yang 0002, Yufan Xu 0001, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
SC | 1 |
| 2025 | Exploiting Dynamic Regular Patterns in Irregular Programs for Efficient VectorizationabstractModern optimizing compilers are able to exploit memory access or computation patterns to generate vectorized codes. However, such patterns in irregular programs are unknown until runtime due to the input dependence. Thus, either compiler’s static optimization or profile-guided optimization cannot represent the patterns for any common input, which leads to suboptimal vectorization. To address the above drawback, we propose DynVec , 1 a framework to automatically exploit regular patterns buried deeply inside irregular programs and apply corresponding optimizations for better vectorization. Due to the integration of workload distribution and the ability to represent instruction features and identify regular patterns with effective feature extraction and data re-arranging methods, DynVec can generate highly efficient vectorized codes for both serial and parallel irregular programs by replacing gather / scatter / reduction operations with optimized operation groups. We evaluate DynVec on optimizing irregular programs such as SpMV and graph programs with representative sparse matrix datasets. The experiment results show that DynVec achieves significant speedup compared to the state-of-the-art implementations across a range of X86 and ARM platforms. Kelun Lei, Shaokang Du, Xin You 0001, Hailong Yang 0002, Zhongzhi Luan, Yi Liu 0013, Depei Qian 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | Efficient Deep Molecular Dynamic Model Training on Heterogeneous SystemabstractMolecular dynamics is a widely adopted simulation method for analyzing the movement of atoms and molecules. Traditional molecular dynamics simulation methods are computationally intensive and difficult to simulate a large number of atoms. In contrast, molecular dynamics based on deep potential models such as DeePMD can leverage deep learning techniques to improve simulation efficiency. Although DeePMD has incorporated mainstream deep learning frameworks, it still suffers from low performance and efficiency during its model training on heterogeneous systems such as CPU and GPU. Particularly, a large number of operators cannot be accelerated by GPU, resulting in low utilization of GPU computational resources. In this paper, we comprehensively analyze the computational bottlenecks and the corresponding root causes of DeePMD. We correspondingly propose several novel optimization strategies. Specifically, for preprocessing, we identify the computation redundancies and the GPU parallelization opportunities for performance optimization. For training, we propose optimization strategies such as operator fusion, redundancy elimination, and concurrent execution of multiple streams and threads in the computation process. Moreover, we apply systematical optimization of computational graphs and operators. The evaluation results show that DeePMD can achieve significant speedups in several cases after applying our proposed optimizations, resulting in a maximum overall speedup of 6.36× with acceptable accuracy. Shaokang Du, Xin You 0001, Hailong Yang 0002, Jing Shang 0001, Zhiwen Xiao, Zhongzhi Luan, Depei Qian 0001 |
ICPADS | 1 |
| 2023 | Accelerating Big Data Application by Eliminating Redundancy on Hadoop ClusterabstractBig data applications are widely adopted to mine valuable information from a tremendous amount of industry data, which is commonly represented as a series of map-reduce operations. Among various map-reduce frameworks, Hadoop is most commonly adopted for data processing at large scale. Although Hadoop eases the development of highly scalable distributed big data applications, inefficient implementation due to poor coding practice and deep software abstractions can cause severe performance issues such as unreasonable slowdown, high response latency, and waste of computing resources, which can lead to unsatisfactory serving delay or significant maintenance cost. In this paper, we first categorize three common types of redundant patterns in big data applications. Then we propose a tool-assisted optimization workflow to detect the redundant patterns automatically, which profiles the application by sampling hardware performance monitoring units. Moreover, we present a profiling visualization method that can help to pinpoint the redundant codes. Based on these approaches, we optimize several big data applications by eliminating redundancies, yielding up to 14.8% performance improvement. Kelun Lei, Shaokang Du, Xin You 0001, Zhibo Xuan, Haoran Kong, Hailong Yang 0002, Jing Shang 0001, Zhiwen Xiao, Zhongzhi Luan, Depei Qian 0001 |
ICPADS | 2 |