EDBT 2026 Demo / reviewers in the wild / expert
Yan Kang 0005
dblp:32/1654-5
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-2673-1947ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CUSPX: Efficient GPU Implementations of Post-Quantum Signature SPHINCS+abstractQuantum computers pose a serious threat to existing cryptographic systems. While Post-Quantum Cryptography (PQC) offers resilience against quantum attacks, its performance limitations often hinder widespread adoption. Among the three National Institute of Standards and Technology (NIST)-selected general-purpose PQC schemes, SPHINCS${}^{+}$is particularly susceptible to these limitations. We introduce CUSPX (CUDASPHINCS${}^{+}$), the first large-scale parallel implementation of SPHINCS${}^{+}$capable of running across 10,000 cores. CUSPX leverages a novel three-level parallelism framework, applying it toalgorithmic parallelism,data parallelism, andhybrid parallelism. Notably, CUSPX introduces parallel Merkle tree construction algorithms for arbitrary parallel scales and several load-balancing solutions, further enhancing performance. By treating tasks parallelism as the top level of parallelism, CUSPX provides a four-level parallel scheme that can run with any number of tasks. Evaluated on a single GeForce RTX 3090 using the SPHINCS${}^{+}$-SHA-256-128s-simple parameter set, CUSPX achieves a single task's signature generation latency of 0.67 ms, demonstrating a 5,105$\times$speedup over a single-thread version and an 18.50$\times$speedup over the previous fastest implementation. Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005, Qiang Wang 0062 |
IEEE Trans. Computers | 4 |
| 2024 | An Example of Parallel Merkle Tree Traversal: Post-Quantum Leighton-Micali Signature on the GPUabstractThe hash-based signature (HBS) is the most conservative and time-consuming among many post-quantum cryptography (PQC) algorithms. Two HBSs, LMS and XMSS, are the only PQC algorithms standardised by the National Institute of Standards and Technology (NIST) now. Existing HBSs are designed based on serial Merkle tree traversal, which is not conducive to taking full advantage of the computing power of parallel architectures such as CPUs and GPUs. We propose a parallel Merkle tree traversal (PMTT), which is tested by implementing LMS on the GPU. This is the first work accelerating LMS on the GPU, which performs well even with over 10,000 cores. Considering different scenarios of algorithmic parallelism and data parallelism, we implement corresponding variants for PMTT. The design of PMTT for algorithmic parallelism mainly considers the execution efficiency of a single task, while that for data parallelism starts with the full utilisation of GPU performance. In addition, we are the first to design a CPU-GPU collaborative processing solution for traversal algorithms to reduce the communication overhead between CPU and GPU. For algorithmic parallelism, our implementation is still 4.48× faster than the ideal time of the state-of-the-art traversal algorithm. For data parallelism, when the number of cores increases from 1 to 8,192, the parallel efficiency is 78.39%. In comparison, our LMS implementation outperforms most existing LMS and XMSS implementations. Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002, Qiang Wang 0062 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Parallel implementations of post-quantum leighton-Micali signature on multiple nodes
Yan Kang 0005, Xiaoshe Dong, Ziheng Wang 0002, Heng Chen 0002, Qiang Wang 0062 |
J. Supercomput. | 1 |
| 2023 | Parallel SHA-256 on SW26010 many-core processor for hashing of multiple messages
Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002 |
J. Supercomput. | 3 |
| 2023 | Efficient GPU Implementations of Post-Quantum Signature XMSSabstractThe National Institute of Standards and Technology (NIST) approved XMSS as part of the post-quantum cryptography (PQC) development effort in 2018. XMSS is currently one of only two standardized PQC algorithms, but its performance limits its use. For example, the fastest record for some standardized parameters still takes more than a minute to generate a keypair. In this article, we present the first GPU implementation for XMSS and its variant XMSS$^{\mathsf {MT}}$. The high parallelism of GPUs is especially effective for reducing latency in key generation and improving throughput for signing and verifying. In order to meet various application scenarios, we provide three parallel XMSS schemes:algorithmic parallelism,multi-keypair data parallelism, andsingle-keypair data parallelism. For these schemes, we design custom parallel strategies that use more than 10,000 cores for all parameters provided by NIST. In addition, we analyze the availability of most previous serial optimizations and explore numerous techniques to fully exploit GPU performance. Our evaluations are made with the XMSSMT-SHA2_20/2_256 parameter set on a GeForce RTX 3090. The result shows the key generation latency is 3.20 ms, a speedup of 21,899× compared to the GPU ported version, which is also 54× speedup faster than the fastest work (174 ms). When 16384 tasks are executed, the throughput (task/s) for signing/verifying in the single-key and multi-key cases is 311,424/415,100 and 145,100/419,887, respectively. Compared to the throughput for signing/verifying (1695/4000) of the fastest work, we obtain a speedup of 184×/104× and 86×/105× in single-key and multi-key cases, respectively. Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Extending τ-Lop to model MPI blocking primitives on shared memory
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Yan Kang 0005, Xingjun Zhang |
J. Supercomput. | 5 |