Qiang Wang 0062

dblp:64/5630-62 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-9179-6611ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Portable GPU Kernel Performance Modeling Method Based on LLVM IR Dynamic Feature Prediction
Qiang Wang 0062, Hao Zheng 0004, Yiru Liu, Chaojun Deng, Ziheng Wang 0002, Xiaoshe Dong
CF1
2026 DP-SWAP: Fast Swapping Strategy Based on Dynamic Programming
Weiduo Chen, Xiaoshe Dong, Qiang Wang 0062
Future Gener. Comput. Syst.3
2026 Mapo: Performance model driven GPU memory access code optimization
Xiaoshe Dong, Junkai Cao, Ruifan Chu, Ziheng Wang 0002, Qiang Wang 0062, Xiuxiu Bai
Future Gener. Comput. Syst.6
2025 ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management
abstract
Due to the limited GPU memory, the performance of large DNNs training is constrained by the unscalable batch size. Existing studies partially address the issue of GPU memory limit through tensor recomputation and swapping, but overlook the exploration of optimal performance. In response, we propose ATP, a recomputation and swapping based GPU memory management framework that aims to maximize training performance by breaking GPU memory constraints. ATP utilizes a throughput model and we propose to evaluate the theoretical peak performance achievable by DNN training on GPU, and provide the optimum memory size required for recomputation and swapping. We optimize the mechanisms for GPU memory pool and CUDA stream control, employ an optimization method to search for specific tensors requiring recomputation and swapping, thereby bringing the actual DNN training performance on ATP closer to theoretical values. Evaluations with different types of large DNN models indicate that ATP achieve throughput improvements ranging from 1.14∼ 1.49×, while support model training exceeding the GPU memory limit by up to 9.2×.
Weiduo Chen, Xiaoshe Dong, Fan Zhang 0139, Bowen Li 0009, Yufei Wang 0008, Qiang Wang 0062
ACM Trans. Archit. Code Optim.6
2025 CUSPX: Efficient GPU Implementations of Post-Quantum Signature SPHINCS+
abstract
Quantum computers pose a serious threat to existing cryptographic systems. While Post-Quantum Cryptography (PQC) offers resilience against quantum attacks, its performance limitations often hinder widespread adoption. Among the three National Institute of Standards and Technology (NIST)-selected general-purpose PQC schemes, SPHINCS${}^{+}$is particularly susceptible to these limitations. We introduce CUSPX (CUDASPHINCS${}^{+}$), the first large-scale parallel implementation of SPHINCS${}^{+}$capable of running across 10,000 cores. CUSPX leverages a novel three-level parallelism framework, applying it toalgorithmic parallelism,data parallelism, andhybrid parallelism. Notably, CUSPX introduces parallel Merkle tree construction algorithms for arbitrary parallel scales and several load-balancing solutions, further enhancing performance. By treating tasks parallelism as the top level of parallelism, CUSPX provides a four-level parallel scheme that can run with any number of tasks. Evaluated on a single GeForce RTX 3090 using the SPHINCS${}^{+}$-SHA-256-128s-simple parameter set, CUSPX achieves a single task's signature generation latency of 0.67 ms, demonstrating a 5,105$\times$speedup over a single-thread version and an 18.50$\times$speedup over the previous fastest implementation.
Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005, Qiang Wang 0062
IEEE Trans. Computers5
2024 pommDNN: Performance optimal GPU memory management for deep neural network training
Weiduo Chen, Xiaoshe Dong, Xinhang Chen, Song Liu 0007, Qin Xia, Qiang Wang 0062
Future Gener. Comput. Syst.6
2024 An Example of Parallel Merkle Tree Traversal: Post-Quantum Leighton-Micali Signature on the GPU
abstract
The hash-based signature (HBS) is the most conservative and time-consuming among many post-quantum cryptography (PQC) algorithms. Two HBSs, LMS and XMSS, are the only PQC algorithms standardised by the National Institute of Standards and Technology (NIST) now. Existing HBSs are designed based on serial Merkle tree traversal, which is not conducive to taking full advantage of the computing power of parallel architectures such as CPUs and GPUs. We propose a parallel Merkle tree traversal (PMTT), which is tested by implementing LMS on the GPU. This is the first work accelerating LMS on the GPU, which performs well even with over 10,000 cores. Considering different scenarios of algorithmic parallelism and data parallelism, we implement corresponding variants for PMTT. The design of PMTT for algorithmic parallelism mainly considers the execution efficiency of a single task, while that for data parallelism starts with the full utilisation of GPU performance. In addition, we are the first to design a CPU-GPU collaborative processing solution for traversal algorithms to reduce the communication overhead between CPU and GPU. For algorithmic parallelism, our implementation is still 4.48× faster than the ideal time of the state-of-the-art traversal algorithm. For data parallelism, when the number of cores increases from 1 to 8,192, the parallel efficiency is 78.39%. In comparison, our LMS implementation outperforms most existing LMS and XMSS implementations.
Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002, Qiang Wang 0062
ACM Trans. Archit. Code Optim.5
2024 Parallel implementations of post-quantum leighton-Micali signature on multiple nodes
Yan Kang 0005, Xiaoshe Dong, Ziheng Wang 0002, Heng Chen 0002, Qiang Wang 0062
J. Supercomput.5
2023 Simplified High Level Parallelism Expression on Heterogeneous Systems through Data Partition Pattern Description
abstract
Abstract With the development of heterogeneous systems, the demand for high-level programming methods that ease heterogeneous programming and produce portable applications has become more urgent. This paper proposes DACL, the data associated computing language. DACL introduces data partition patterns to achieve architecture-independent parallelism expression. Meanwhile, DACL provides simplified language extensions, as well as programming features such as serialization of the computing process, parameterization of data attributes and modularity, thus reducing the difficulty of heterogeneous programming and improving programming productivity. The operational semantics show that DACL enables different levels of parallelism degree calculation and retains data access patterns, reserving optimization potential. To support cross-platform execution, the currently implemented source-to-source compilers employ OpenMP and OpenCL as the backend. We reconstructed multiple benchmarks selected from the Parboil and Rodinia benchmark suits with DACL and conducted a comparison test on CPU, GPU and MIC platforms. The code size of each rebuilt benchmark is roughly equivalent to that of the serial code, which is only 13%–64% of the benchmark OpenCL code. With the support of the compilation system, the reconstructed code can execute on different processors without modification, yielding a competitive or better performance to that of the manually written benchmark code.
Shusen Wu, Xiaoshe Dong, Heng Chen 0002, Qiang Wang 0062, Zhengdong Zhu
Comput. J.5
2022 Status, challenges and trends of data-intensive supercomputing
Jia Wei 0002, Pei Ren, Yujia Lei, Yuqi Qu, Qiyu Jiang, Xiaoshe Dong, Weiguo Wu, Qiang Wang 0062, Xingjun Zhang
CCF Trans. High Perform. Comput.10
2021 Performance evaluation of convolutional neural network on Tianhe-3 prototype
Weiduo Chen, Xiaoshe Dong, Heng Chen 0002, Qiang Wang 0062, Xingda Yu, Xingjun Zhang
J. Supercomput.4