VLDB 2026 Research / reviewers in the wild / expert
Heng Chen 0002
dblp:00/890-2
· DBLP profile ↗
21ranked-venue papers
1as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CUSPX: Efficient GPU Implementations of Post-Quantum Signature SPHINCS+abstractQuantum computers pose a serious threat to existing cryptographic systems. While Post-Quantum Cryptography (PQC) offers resilience against quantum attacks, its performance limitations often hinder widespread adoption. Among the three National Institute of Standards and Technology (NIST)-selected general-purpose PQC schemes, SPHINCS${}^{+}$is particularly susceptible to these limitations. We introduce CUSPX (CUDASPHINCS${}^{+}$), the first large-scale parallel implementation of SPHINCS${}^{+}$capable of running across 10,000 cores. CUSPX leverages a novel three-level parallelism framework, applying it toalgorithmic parallelism,data parallelism, andhybrid parallelism. Notably, CUSPX introduces parallel Merkle tree construction algorithms for arbitrary parallel scales and several load-balancing solutions, further enhancing performance. By treating tasks parallelism as the top level of parallelism, CUSPX provides a four-level parallel scheme that can run with any number of tasks. Evaluated on a single GeForce RTX 3090 using the SPHINCS${}^{+}$-SHA-256-128s-simple parameter set, CUSPX achieves a single task's signature generation latency of 0.67 ms, demonstrating a 5,105$\times$speedup over a single-thread version and an 18.50$\times$speedup over the previous fastest implementation. Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005, Qiang Wang 0062 |
IEEE Trans. Computers | 3 |
| 2024 | An Example of Parallel Merkle Tree Traversal: Post-Quantum Leighton-Micali Signature on the GPUabstractThe hash-based signature (HBS) is the most conservative and time-consuming among many post-quantum cryptography (PQC) algorithms. Two HBSs, LMS and XMSS, are the only PQC algorithms standardised by the National Institute of Standards and Technology (NIST) now. Existing HBSs are designed based on serial Merkle tree traversal, which is not conducive to taking full advantage of the computing power of parallel architectures such as CPUs and GPUs. We propose a parallel Merkle tree traversal (PMTT), which is tested by implementing LMS on the GPU. This is the first work accelerating LMS on the GPU, which performs well even with over 10,000 cores. Considering different scenarios of algorithmic parallelism and data parallelism, we implement corresponding variants for PMTT. The design of PMTT for algorithmic parallelism mainly considers the execution efficiency of a single task, while that for data parallelism starts with the full utilisation of GPU performance. In addition, we are the first to design a CPU-GPU collaborative processing solution for traversal algorithms to reduce the communication overhead between CPU and GPU. For algorithmic parallelism, our implementation is still 4.48× faster than the ideal time of the state-of-the-art traversal algorithm. For data parallelism, when the number of cores increases from 1 to 8,192, the parallel efficiency is 78.39%. In comparison, our LMS implementation outperforms most existing LMS and XMSS implementations. Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002, Qiang Wang 0062 |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | Parallel implementations of post-quantum leighton-Micali signature on multiple nodes
Yan Kang 0005, Xiaoshe Dong, Ziheng Wang 0002, Heng Chen 0002, Qiang Wang 0062 |
J. Supercomput. | 4 |
| 2023 | Simplified High Level Parallelism Expression on Heterogeneous Systems through Data Partition Pattern DescriptionabstractAbstract With the development of heterogeneous systems, the demand for high-level programming methods that ease heterogeneous programming and produce portable applications has become more urgent. This paper proposes DACL, the data associated computing language. DACL introduces data partition patterns to achieve architecture-independent parallelism expression. Meanwhile, DACL provides simplified language extensions, as well as programming features such as serialization of the computing process, parameterization of data attributes and modularity, thus reducing the difficulty of heterogeneous programming and improving programming productivity. The operational semantics show that DACL enables different levels of parallelism degree calculation and retains data access patterns, reserving optimization potential. To support cross-platform execution, the currently implemented source-to-source compilers employ OpenMP and OpenCL as the backend. We reconstructed multiple benchmarks selected from the Parboil and Rodinia benchmark suits with DACL and conducted a comparison test on CPU, GPU and MIC platforms. The code size of each rebuilt benchmark is roughly equivalent to that of the serial code, which is only 13%–64% of the benchmark OpenCL code. With the support of the compilation system, the reconstructed code can execute on different processors without modification, yielding a competitive or better performance to that of the manually written benchmark code. Shusen Wu, Xiaoshe Dong, Heng Chen 0002, Qiang Wang 0062, Zhengdong Zhu |
Comput. J. | 3 |
| 2023 | Parallel SHA-256 on SW26010 many-core processor for hashing of multiple messages
Ziheng Wang 0002, Xiaoshe Dong, Yan Kang 0005, Heng Chen 0002 |
J. Supercomput. | 4 |
| 2023 | Efficient GPU Implementations of Post-Quantum Signature XMSSabstractThe National Institute of Standards and Technology (NIST) approved XMSS as part of the post-quantum cryptography (PQC) development effort in 2018. XMSS is currently one of only two standardized PQC algorithms, but its performance limits its use. For example, the fastest record for some standardized parameters still takes more than a minute to generate a keypair. In this article, we present the first GPU implementation for XMSS and its variant XMSS$^{\mathsf {MT}}$. The high parallelism of GPUs is especially effective for reducing latency in key generation and improving throughput for signing and verifying. In order to meet various application scenarios, we provide three parallel XMSS schemes:algorithmic parallelism,multi-keypair data parallelism, andsingle-keypair data parallelism. For these schemes, we design custom parallel strategies that use more than 10,000 cores for all parameters provided by NIST. In addition, we analyze the availability of most previous serial optimizations and explore numerous techniques to fully exploit GPU performance. Our evaluations are made with the XMSSMT-SHA2_20/2_256 parameter set on a GeForce RTX 3090. The result shows the key generation latency is 3.20 ms, a speedup of 21,899× compared to the GPU ported version, which is also 54× speedup faster than the fastest work (174 ms). When 16384 tasks are executed, the throughput (task/s) for signing/verifying in the single-key and multi-key cases is 311,424/415,100 and 145,100/419,887, respectively. Compared to the throughput for signing/verifying (1695/4000) of the fastest work, we obtain a speedup of 184×/104× and 86×/105× in single-key and multi-key cases, respectively. Ziheng Wang 0002, Xiaoshe Dong, Heng Chen 0002, Yan Kang 0005 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Flexible Supervision System: A Fast Fault-Tolerance Strategy for Cloud Applications in Cloud-Edge Collaborative Environments
Weilin Cai, Heng Chen 0002, Zhimin Zhuo, Ziheng Wang 0002, Ninggang An |
NPC | 2 |
| 2022 | LogSC: Model-based one-sided communication performance estimation
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Xingjun Zhang |
Future Gener. Comput. Syst. | 2 |
| 2022 | C-Lop: Accurate contention-based modeling of MPI concurrent communication
Ziheng Wang 0002, Heng Chen 0002, Weiling Cai, Xiaoshe Dong, Xingjun Zhang |
Parallel Comput. | 2 |
| 2022 | Implementation and optimization of ChaCha20 stream cipher on sunway taihuLight supercomputer
Weilin Cai, Heng Chen 0002, Ziheng Wang 0002, Xingjun Zhang |
J. Supercomput. | 2 |
| 2022 | SunwayURANS: 3D full-annulus URANS simulations of transonic axial compressors on Sunway TaihuLight
Heng Chen 0002, Ziheng Wang 0002, Xiaoshe Dong, Xingjun Zhang |
J. Supercomput. | 1 |
| 2022 | Thou code: a triple-erasure-correcting horizontal code with optimal update complexity
Ningjing Liang, Xingjun Zhang, Heng Chen 0002, Changjiang Zhang |
J. Supercomput. | 3 |
| 2022 | Extending τ-Lop to model MPI blocking primitives on shared memory
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Yan Kang 0005, Xingjun Zhang |
J. Supercomput. | 2 |
| 2021 | Performance evaluation of convolutional neural network on Tianhe-3 prototype
Weiduo Chen, Xiaoshe Dong, Heng Chen 0002, Qiang Wang 0062, Xingda Yu, Xingjun Zhang |
J. Supercomput. | 3 |
| 2020 | H2Pregel : A partition-based hybrid hierarchical graph computation approach
Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
Future Gener. Comput. Syst. | 3 |
| 2020 | Fine-grained scheduling in multi-resource clusters
Mosong Zhou, Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
J. Supercomput. | 3 |
| 2018 | A Runtime Available Resource Capacity Evaluation Model Based on the Concept of Similar TasksabstractA mismatch between resource supply and demand in cloud computing leads to inefficient utilization of resources or performance degradation. Therefore, this paper establishes a runtime model to evaluate the available capacity of computing resources on the basis of similar tasks. This model takes advantage of a characteristic of cloud workload; that is, similar tasks in cloud computing have a similar execution logic. The model evaluates the available resource capacity according to task similarity, thus avoiding any impact on the resource consumption of existing benchmarks. We apply the model to propose a resource capacity evaluation method called Caipan, which considers numerous factors according to resource type. This method obtains accurate results in a timely manner at little cost. We use the results of Caipan to develop some algorithms that aim to match resource supply and demand, and improve cloud platform performance. We test the Caipan method and the Caipan-based algorithms in both dedicated and real-world cloud environments. The test results show that the Caipan method obtains the available resource capacity both accurately and in a timely manner, and effectively supports the optimization of both algorithms and platforms. Moreover, algorithms based on Caipan reduce the mismatch between resource supply and demand, and significantly improve cloud platform performance. Mosong Zhou, Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
Comput. J. | 3 |
| 2018 | IncPregel: an incremental graph parallel computation model
Xiaoshe Dong, Heng Chen 0002, Yinfeng Wang |
Frontiers Comput. Sci. | 3 |
| 2017 | Small files storing and computing optimization in Hadoop parallel renderingabstractSummary Hadoop framework has been widely used in animation industries to build a large scale, high performance parallel rendering system. However, Hadoop Distributed File System (HDFS) and the MapReduce programming model are designed to manage large files and suffer performance penalty while rendering and storing small files in a rendering system. Therefore, a method that merges small files based on two intelligent algorithms is proposed to solve the problem. The method uses Particle Swarm Optimization (PSO) to select the optimal merge values for multiple sets of scenes and then uses Support Vector Machine (SVM) to generate a general SVM model which can be used to get the optimal merge value for any scene, by mainly considering the rendering time, memory limitation and other indicators. Then, the method takes advantage of frame‐to‐frame coherence to merge files in the same scene in an interval‐based way with the optimal merge value. Finally, the proposed method is compared with the naive method under three different render scenes. Experimental results show that the proposed method significantly reduces the number of small files and render tasks, and improves the storage efficiency and computing efficiency. Copyright © 2016 John Wiley & Sons, Ltd. Yizhi Zhang, Zhengdong Zhu, Honglin Cui, Xiaoshe Dong, Heng Chen 0002 |
Concurr. Comput. Pract. Exp. | 5 |
| 2015 | Thread Count Prediction Model: Dynamically Adjusting Threads for Heterogeneous Many-Core SystemsabstractDetermining an appropriate thread count for a multithread application running on a heterogeneous many-core system is crucial for improving computing performance and reducing energy consumption. This paper investigates the interrelation between thread count and computing performance of applications, and designs a prediction model of the optimum thread count on the basis of Amdahl's law combined with regression analysis theory to improve computing performance and reduce energy consumption. The prediction model can estimate the optimum tread count relying on the program running behaviors and the architecture characteristics of heterogeneous many-core system. Using the estimated optimum thread count, the number of the active hardware threads and processing cores on the many-core processor is dynamically adjusted in the process of thread mapping to improve the energy efficiency of entire heterogeneous many-core system. The experimental results show that, using this paper proposed thread count prediction model, on an average, the computing performance is improved by 48.6%, energy consumption is reduced by 59%, and additional overhead introduced is 2.03% compared with that of the traditional thread mapping for the PARSEC benchmark programs run on an Intel MIC heterogeneous many-core system. Tao Ju 0002, Weiguo Wu, Heng Chen 0002, Zhengdong Zhu, Xiaoshe Dong |
ICPADS | 3 |
| 2008 | EntityTrust: Feedback credibility-based global reputation mechanism in cooperative computing systemabstractTrust and reputation are important decision-making factors in cooperative computing systems. It is a fundamental but challenging task for reputation systems to estimate entity's reputation accurately and efficiently in distributed environment. We propose EntityTrust, a global reputation mechanism in cooperative computing systems. EntityTrust introduces a new direct feedback metrics to reflect the dynamic feature of trust. Besides that, an entity's feedback credibility can affect its reputation in a direct way. Experimental results show that EntityTrust can improve the accuracy and efficiency of evaluation for global reputation. Moreover, it can effectively combat malicious behaviors presented in the cooperative communities. Yiduo Mei, Xiaoshe Dong, Zhenhua Tian, Shangyuan Guan, Heng Chen 0002 |
CSCWD | 5 |