VLDB 2026 Research / reviewers in the wild / expert
Hao Yang 0062
dblp:54/4089-62
· DBLP profile ↗
14ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-9735-255XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 5 · 4 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration
Hanyu Wei, Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Yunlei Zhao |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Optimized Implementation of NTRU on RISC-V Platform
Lu Zhou 0002, Hao Yang 0062, Zhe Liu 0001 |
ProvSec | 3 |
| 2025 | CBPSPX: A CUDA-Based Batch Parallel Optimization of Post-Quantum Signature SPHINCS+abstractSecurity and privacy are critical in cloud-based Internet of Things (IoT) and Artificial Intelligence of Things (AIoT) applications. As quantum computing advances, Post-Quantum Cryptography (PQC) has emerged as a key technology for ensuring security in future IoT and AIoT architectures. SPHINCS+, a leading post-quantum signature algorithm, has been selected by the National Institute of Standards and Technology (NIST) as one of the next-generation signature standards. However, due to its complex structure and extensive hash operations, SPHINCS+ suffers from slower signature generation and verification compared to other post-quantum algorithms. Consequently, accelerating SPHINCS+ is essential for adapting it to IoT environments. This paper presents a CUDA-based Batch Parallel optimization of SPHINCS+ (CBPSPX), which fully utilizes the computing resources of NVIDIA Graphics Processing Units (GPUs) to enhance the performance of SPHINCS+. Specifically, we propose the Thread Utilization Efficiency Index (TUEI), which can be used to theoretically evaluate the effectiveness of various parallel methods. Then, we propose an intra-block batch processing model that dynamically adjusts parallel task scales within a block to optimize throughput, making it particularly suitable for IoT scenarios requiring high-throughput large-scale device authentication. Meanwhile, we divide the signature generation process into three sub-processes and adopt different parallel strategies based on the thread requirements of each sub-process to maximize the value of TUEI. For the signature verification process, we propose a columnar storage strategy to replace the traditional row storage structure, which significantly improves the performance of batch signature verification. Experimental results indicate that our SPHINCS+ implementations across all three parameter sets are better than previous optimized GPU-based implementations and achieve speedups of 1.4× to 2.5× for signature generation and 4.6× to 11.3× for signature verification on GPU RTX 3090. Jiafei Wu, Hao Yang 0062, Zhe Liu 0001 |
IEEE Internet Things J. | 4 |
| 2025 | cuML-DSA: Optimized Signing Procedure and Server-Oriented GPU Design for ML-DSAabstractThe threat posed by quantum computing has precipitated an urgent need for post-quantum cryptography. Recently, the post-quantum digital signature draft FIPS 204 has been published, delineating the details of the ML-DSA, which is derived from the CRYSTALS-Dilithium. Despite these advancements, server environments, especially those equipped with GPU devices necessitating high-throughput signing, remain entrenched in classical schemes. A conspicuous void exists in the realm of GPU implementation or server-specific designs for ML-DSA. In this paper, we propose the first server-oriented GPU design tailored for the ML-DSA signing procedure in high-throughput servers. We introduce several innovative theoretical optimizations to bolster performance, including depth-prior sparse ternary polynomial multiplication, the branch elimination method, and the rejection-prioritized checking order. Furthermore, exploiting server-oriented features, we propose a comprehensive GPU hardware design, augmented by a suite of GPU implementation optimizations to further amplify performance. Additionally, we present variants for sampling sparse polynomials, thereby streamlining our design. The deployment of our implementation on both server-grade and commercial GPUs demonstrates significant speedups, ranging from 284.9× to 485.3× against the CPU baseline, and an improvement of up to 60.9% compared to related work, affirming the effectiveness and efficiency of the proposed GPU architecture for ML-DSA signing procedure. Shiyu Shen 0001, Hao Yang 0062, Yunlei Zhao |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | Noya: An Efficient, Flexible and Secure CNN Inference Model Based on Homomorphic Encryption
Fengyuan Qiu, Hao Yang 0062, Lu Zhou 0002, Zhe Liu 0001 |
SecureComm (1) | 2 |
| 2024 | BCPIR: More Efficient Keyword PIR via Block Building Codewords
Shuquan Wang, Hao Yang 0062, Lu Zhou 0002 |
SecureComm (1) | 2 |
| 2024 | Leveraging GPU in Homomorphic Encryption: Framework Design and Analysis of BFV VariantsabstractHomomorphic Encryption (HE) enhances data security by enabling computations on encrypted data, advancing privacy-focused computations. The BFV scheme, a promising HE scheme, raises considerable performance challenges. Graphics Processing Units (GPUs), with considerable parallel processing abilities, offer an effective solution. In this work, we present an in-depth study on accelerating and comparing BFV variants on GPUs, including Bajard-Eynard-Hasan-Zucca (BEHZ), Halevi-Polyakov-Shoup (HPS), and recent variants. We introduce a universal framework for all variants, propose optimized BEHZ implementation, and first support HPS variants with large parameter sets on GPUs. We also optimize low-level arithmetic and high-level operations, minimizing instructions for modular operations, enhancing hardware utilization for base conversion, and implementing efficient reuse strategies and fusion methods to reduce computational and memory consumption. Leveraging our framework, we offer comprehensive comparative analyses. Performance evaluation shows a 31.9$\times$speedup over OpenFHE running on a multi-threaded CPU and 39.7% and 29.9% improvement for tensoring and relinearization over the state-of-the-art GPU BEHZ implementation. The leveled HPS variant records up to 4$\times$speedup over other variants, positioning it as a highly promising alternative for specific applications. Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Lu Zhou 0002, Zhe Liu 0001, Yunlei Zhao |
IEEE Trans. Computers | 2 |
| 2024 | Phantom: A CUDA-Accelerated Word-Wise Homomorphic Encryption LibraryabstractHomomorphic encryption (HE) is a promising technique for privacy-preserving computations, especially the word-wise HE schemes that allow batching. However, the high computational overhead hinders the deployment of HE in real-word applications. GPUs are often used to accelerate execution, but a comprehensive performance comparison of different schemes on the same platform is still missing. In this work, we fill this gap by implementing three word-wise HE schemes BGV, BFV, and CKKS on GPU, with both theoretical and engineering optimizations. We enhance the hybrid key-switching technique, significantly reducing the computational and memory overhead. We explore several kernel fusing strategies to reuse data, resulting in reduced memory access and IO latency, and enhancing the overall performance. By comparing with the state-of-the-art works, we demonstrate the effectiveness of our implementation. Meanwhile, we introduce a unified framework that finely integrates our implementation of the three schemes, covering almost all scheme functions and homomorphic operations. We optimize the management of pre-computation, RNS bases, and memory in the framework, to provide efficient and low-latency data access and transfer. Based on this framework, we provide a thorough benchmark of the three schemes, which can serve as a reference for scheme selection and implementation in constructing privacy-preserving applications. Hao Yang 0062, Shiyu Shen 0001, Wangchen Dai, Lu Zhou 0002, Zhe Liu 0001, Yunlei Zhao |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | cuXCMP: CUDA-Accelerated Private Comparison Based on Homomorphic EncryptionabstractPrivate comparison schemes constructed on homomorphic encryption offer the noninteractive and parallelizable features, and have advantages in communication bandwidth and performance. In this work, we propose cuXCMP, an extension of the privacy comparison scheme XCMP (AsiaCCS 2018). We address the relatively small input domain and the incompletely expressible output of XCMP, by modifying the encoding method and devising a constant term extraction (CTX) approach. Then, we describe a method for constructing privacy-preserving decision tree (PPDT) using this scheme. Considering the high computational overhead of CTX, we exploit the massive parallelism of the GPU to accelerate this function. Based on the results of the kernel profiling, we utilize several optimization techniques to improve the performance, including using multiple CUDA streams, reducing the grid dimension, kernel fusion, etc. By accelerating this function, we boost the execution time of the scheme and demonstrate 130× and 1.9× speedups for CTX and cuXCMP, respectively, as well as a 35% reduction in the evaluation time of PPDT. Hao Yang 0062, Shiyu Shen 0001, Zhe Liu 0001, Yunlei Zhao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | High-Throughput GPU Implementation of Dilithium Post-Quantum Digital SignatureabstractDigital signatures are fundamental building blocks in various protocols to provide integrity and authenticity. The development of the quantum computing has raised concerns about the security guarantees afforded by classical signature schemes. CRYSTALS-Dilithium is an efficient post-quantum digital signature scheme based on lattice cryptography and has been selected as the primary algorithm for standardization by the National Institute of Standards and Technology. In this work, we present a high-throughput GPU implementation of Dilithium. For individual operations, we employ a range of computational and memory optimizations to overcome sequential constraints, reduce memory usage and IO latency, address bank conflicts, and mitigate pipeline stalls. This results in high and balanced compute throughput and memory throughput for each operation. In terms of concurrent task processing, we leverage task-level batching to fully utilize parallelism and implement a memory pool mechanism for rapid memory access. We propose a dynamic task scheduling mechanism to improve multiprocessor occupancy and significantly reduce execution time. Furthermore, we apply asynchronous computing and launch multiple streams to hide data transfer latencies and maximize the computing capabilities of both CPU and GPU. Across all three security levels, our GPU implementation achieves over 160× speedups for signing and over 80× speedups for verification on both commercial and server-grade GPUs. This achieves microsecond-level amortized execution times for each task, offering a high-throughput and quantum-resistant solution suitable for a wide array of applications in real systems. Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Zhe Liu 0001, Yunlei Zhao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Poster Abstract: CNN-guardian: Secure Neural Network Inference Acceleration on Edge GPUabstractThe rapid development of AI applications powered by deep learning in edge devices boosts the opportunity for real-time health monitoring. To address the potential privacy concern in the inference phase, homomorphic encryption (HE) is an alternative solution that encrypts inference data without exposing raw data and has several distinct advantages, (i.e., single-round communication, lightweight bandwidth consumption, and non-interactive computation). However, the computational overhead on the current HE-based privacy-preserving inference necessitates a substantial amount of time, which is not feasible for some real-time applications on edge devices. To address this issue, we propose CNN-guardian, a unified and compact neural network structure for real-time inference in HE-based inference on edge GPU. CNN-guardian designs a HE-friendly neural network and GPU engine that optimizes HE operations to accelerate the inference in the HE domain. Qipeng Xie, Hao Yang 0062, Linshan Jiang, Siyang Jiang, Shiyu Shen 0001, Salabat Khan, Zhe Liu 0001, Kaishun Wu |
SenSys | 2 |
| 2023 | CARM: CUDA-Accelerated RNS Multiplication in Word-Wise Homomorphic Encryption Schemes for Internet of ThingsabstractHomomorphic encryption (HE), which allows computation over encrypted data, has often been used to preserve privacy. However, the computationally heavy nature and complexity of network topologies make the deployment of HE schemes in the Internet of Things (IoT) scenario difficult. In this work, we propose CARM, the first optimized GPU implementation that covers BGV, BFV and CKKS, targeting for accelerating homomorphic multiplication using GPU in heterogeneous IoT systems. Our solution is suitable for accelerating RNS homomorphic multiplication on both high-performance and embedded GPUs, as it is a parametric and generic design and offers various trade-offs between resource and efficiency. We offer constant-time low-level arithmetic with minimum instructions and memory usage, as well as performance- and memory-prior configurations. Through this, we can provide more real-time evaluation results and relieve the computational pressure on cloud devices. We deploy our implementations on two GPUs. Compared to the CPU implementation, we achieve up to$378.4\times$,$234.5\times$, and$287.2\times$speedup for homomorphic multiplication of BGV, BFV, and CKKS on Tesla V100S, and$8.8\times$,$9.2\times$, and$10.3\times$on Jetson AGX Xavier, respectively. Shiyu Shen 0001, Hao Yang 0062, Zhe Liu 0001, Yunlei Zhao |
IEEE Trans. Computers | 2 |
| 2022 | Privacy Preserving Federated Learning Using CKKS Homomorphic Encryption
Fengyuan Qiu, Hao Yang 0062, Lu Zhou 0002, Chuan Ma 0001, Liming Fang 0001 |
WASA (1) | 2 |
| 2020 | An Efficient and Scalable Sparse Polynomial Multiplication Accelerator for LAC on FPGAabstractLAC, a Ring-LWE based scheme, has shortlisted for the second round evaluation of the National Institute of Standards and Technology Post-Quantum Cryptography (NIST-PQC) Standardization. FPGAs are widely used to design accelerators for cryptographic schemes, especially in resource-constrained scenarios, such as IoT. Sparse Polynomial Multiplication (SPM) is the most compute-intensive routine in LAC. Designing an accelerator for SPM on FPGA can significantly improve the performance of LAC. However, as far as we know, there are currently no works related to the hardware implementation of SPM for LAC. In this paper, the proposed efficient and scalable SPM accelerator fills this gap. More concretely, we firstly develop the Dual-For-Loop-Parallel (DFLP) technique to optimize the accelerator's parallel design. This technique can achieve 2x performance improvement compared with the previous works. Secondly, we design a hardware-friendly modular reduction algorithm for the modulus 251. Our method not only saves hardware resources but also improves performance. Then, we launch a detailed analysis and optimization of the pipeline design, achieving a frequency improvement of up to 34%. Finally, our design is scalable, and we can achieve various performance-area trade-offs through parameter p. Our results demonstrate that the proposed design can achieve a very considerable performance improvement with moderate hardware area costs. For example, our medium-scale architecture for LAC-128 takes only 783 LUTs, 432 FFs, 5BRAMs, and no DSP on an Artix-7 FPGA and can complete LAC's polynomial multiplication in 8512 cycles at a frequency of 202MHz. Jipeng Zhang 0001, Zhe Liu 0001, Hao Yang 0062, Junhao Huang 0001, Weibin Wu 0003 |
ICPADS | 3 |