Ziheng Wang 0002

dblp:79/10743-2 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0001-5064-2376ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 8 first-author · 14 since 2021
YearPublicationVenuePosition
2026 A Portable GPU Kernel Performance Modeling Method Based on LLVM IR Dynamic Feature Prediction
Qiang Wang 0062, Hao Zheng 0004, Yiru Liu, Chaojun Deng, Ziheng Wang 0002, Xiaoshe Dong
CF7
2026 Mapo: Performance model driven GPU memory access code optimization
Xiaoshe Dong, Junkai Cao, Ruifan Chu, Ziheng Wang 0002, Qiang Wang 0062, Xiuxiu Bai
Future Gener. Comput. Syst.5
2026 BLG-Tuning: Benchmark-Based Low-Cost General-Purpose I/O Modeling and Tuning
abstract
I/O performance has become a major bottleneck for many data-intensive applications. Each layer of the parallel I/O stack provides parameters that can optimize I/O performance, but determining the optimal performance parameters based on the operating configuration is a challenge. Previous work has required separate performance models for different programs for tuning, which is very costly in term of measurement data. We propose BLG-Tuning: a B enchmark-based L ow-cost G eneral-purpose I/O Modeling and Tuning. BLG-Tuning maps application I/O loads to benchmark parameters and uses the benchmark-trained performance model to achieve I/O performance prediction and thus avoid the additional computing and communication overhead for measurement. For applications, BLG-Tuning collects the application characteristics to calibrate the performance model and improve prediction accuracy. Experience shows that BLG-Tuning predicts the I/O time of MADbench2, Flash-IO, S3D-IO, BT-IO, and LAMMPS with MAPE of 22.2%, 18.2%, 29.3%, 21.5%, and 34.8%, respectively. After tuning, the five applications obtain I/O speedup from 5.6× to 27.3×.
Ziheng Wang 0002, Yuchao Wu, Xiaoshe Dong
ACM Trans. Archit. Code Optim.1
2025 CUSPX: Efficient GPU Implementations of Post-Quantum Signature SPHINCS+
abstract
Quantum computers pose a serious threat to existing cryptographic systems. While Post-Quantum Cryptography (PQC) offers resilience against quantum attacks, its performance limitations often hinder widespread adoption. Among the three National Institute of Standards and Technology (NIST)-selected general-purpose PQC schemes, SPHINCS${}^{+}$is particularly susceptible to these limitations. We introduce CUSPX (CUDASPHINCS${}^{+}$), the first large-scale parallel implementation of SPHINCS${}^{+}$capable of running across 10,000 cores. CUSPX leverages a novel three-level parallelism framework, applying it toalgorithmic parallelism,data parallelism, andhybrid parallelism. Notably, CUSPX introduces parallel Merkle tree construction algorithms for arbitrary parallel scales and several load-balancing solutions, further enhancing performance. By treating tasks parallelism as the top level of parallelism, CUSPX provides a four-level parallel scheme that can run with any number of tasks. Evaluated on a single GeForce RTX 3090 using the SPHINCS${}^{+}$-SHA-256-128s-simple parameter set, CUSPX achieves a single task's signature generation latency of 0.67 ms, demonstrating a 5,105$\times$speedup over a single-thread version and an 18.50$\times$speedup over the previous fastest implementation.
Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005, Qiang Wang 0062
IEEE Trans. Computers1
2024 An Example of Parallel Merkle Tree Traversal: Post-Quantum Leighton-Micali Signature on the GPU
abstract
The hash-based signature (HBS) is the most conservative and time-consuming among many post-quantum cryptography (PQC) algorithms. Two HBSs, LMS and XMSS, are the only PQC algorithms standardised by the National Institute of Standards and Technology (NIST) now. Existing HBSs are designed based on serial Merkle tree traversal, which is not conducive to taking full advantage of the computing power of parallel architectures such as CPUs and GPUs. We propose a parallel Merkle tree traversal (PMTT), which is tested by implementing LMS on the GPU. This is the first work accelerating LMS on the GPU, which performs well even with over 10,000 cores. Considering different scenarios of algorithmic parallelism and data parallelism, we implement corresponding variants for PMTT. The design of PMTT for algorithmic parallelism mainly considers the execution efficiency of a single task, while that for data parallelism starts with the full utilisation of GPU performance. In addition, we are the first to design a CPU-GPU collaborative processing solution for traversal algorithms to reduce the communication overhead between CPU and GPU. For algorithmic parallelism, our implementation is still 4.48× faster than the ideal time of the state-of-the-art traversal algorithm. For data parallelism, when the number of cores increases from 1 to 8,192, the parallel efficiency is 78.39%. In comparison, our LMS implementation outperforms most existing LMS and XMSS implementations.
Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002, Qiang Wang 0062
ACM Trans. Archit. Code Optim.1
2024 Parallel implementations of post-quantum leighton-Micali signature on multiple nodes
Yan Kang 0005, Xiaoshe Dong, Ziheng Wang 0002, Heng Chen 0002, Qiang Wang 0062
J. Supercomput.3
2023 Parallel SHA-256 on SW26010 many-core processor for hashing of multiple messages
Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002
J. Supercomput.1
2023 Efficient GPU Implementations of Post-Quantum Signature XMSS
abstract
The National Institute of Standards and Technology (NIST) approved XMSS as part of the post-quantum cryptography (PQC) development effort in 2018. XMSS is currently one of only two standardized PQC algorithms, but its performance limits its use. For example, the fastest record for some standardized parameters still takes more than a minute to generate a keypair. In this article, we present the first GPU implementation for XMSS and its variant XMSS$^{\mathsf {MT}}$. The high parallelism of GPUs is especially effective for reducing latency in key generation and improving throughput for signing and verifying. In order to meet various application scenarios, we provide three parallel XMSS schemes:algorithmic parallelism,multi-keypair data parallelism, andsingle-keypair data parallelism. For these schemes, we design custom parallel strategies that use more than 10,000 cores for all parameters provided by NIST. In addition, we analyze the availability of most previous serial optimizations and explore numerous techniques to fully exploit GPU performance. Our evaluations are made with the XMSSMT-SHA2_20/2_256 parameter set on a GeForce RTX 3090. The result shows the key generation latency is 3.20 ms, a speedup of 21,899× compared to the GPU ported version, which is also 54× speedup faster than the fastest work (174 ms). When 16384 tasks are executed, the throughput (task/s) for signing/verifying in the single-key and multi-key cases is 311,424/415,100 and 145,100/419,887, respectively. Compared to the throughput for signing/verifying (1695/4000) of the fastest work, we obtain a speedup of 184×/104× and 86×/105× in single-key and multi-key cases, respectively.
Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005
IEEE Trans. Parallel Distributed Syst.1
2022 Flexible Supervision System: A Fast Fault-Tolerance Strategy for Cloud Applications in Cloud-Edge Collaborative Environments
Weilin Cai, Heng Chen 0002, Zhimin Zhuo, Ziheng Wang 0002, Ninggang An
NPC4
2022 LogSC: Model-based one-sided communication performance estimation
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Xingjun Zhang
Future Gener. Comput. Syst.1
2022 C-Lop: Accurate contention-based modeling of MPI concurrent communication
Ziheng Wang 0002, Heng Chen 0002, Weiling Cai, Xiaoshe Dong, Xingjun Zhang
Parallel Comput.1
2022 Implementation and optimization of ChaCha20 stream cipher on sunway taihuLight supercomputer
Weilin Cai, Heng Chen 0002, Ziheng Wang 0002, Xingjun Zhang
J. Supercomput.3
2022 SunwayURANS: 3D full-annulus URANS simulations of transonic axial compressors on Sunway TaihuLight
Heng Chen 0002, Ziheng Wang 0002, Xiaoshe Dong, Xingjun Zhang
J. Supercomput.2
2022 Extending τ-Lop to model MPI blocking primitives on shared memory
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Yan Kang 0005, Xingjun Zhang
J. Supercomput.1