Shiyu Shen 0001

dblp:288/3142-1 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-7287-4223ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 BW-KEM: Robust and Versatile MLWE-Based KEM with Barnes-Wall Lattices
Hengchuan Zou, Shiyu Shen 0001, Yunlei Zhao
PKC (1)3
2026 cuFalcon: An Adaptive Parallel GPU Implementation for High-Performance Falcon Acceleration
Hanyu Wei, Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Yunlei Zhao
IEEE Trans. Parallel Distributed Syst.3
2025 OSKR/OKAI: Systematic Optimization of Key Encapsulation Mechanisms from Module Lattice
Shiyu Shen 0001, Zhichuang Liang, Jieyu Zheng, Hanyu Wei, Yang Wang 0050, Zhenfeng Zhang, Yunlei Zhao
J. Comput. Sci. Technol.1
2025 cuML-DSA: Optimized Signing Procedure and Server-Oriented GPU Design for ML-DSA
abstract
The threat posed by quantum computing has precipitated an urgent need for post-quantum cryptography. Recently, the post-quantum digital signature draft FIPS 204 has been published, delineating the details of the ML-DSA, which is derived from the CRYSTALS-Dilithium. Despite these advancements, server environments, especially those equipped with GPU devices necessitating high-throughput signing, remain entrenched in classical schemes. A conspicuous void exists in the realm of GPU implementation or server-specific designs for ML-DSA. In this paper, we propose the first server-oriented GPU design tailored for the ML-DSA signing procedure in high-throughput servers. We introduce several innovative theoretical optimizations to bolster performance, including depth-prior sparse ternary polynomial multiplication, the branch elimination method, and the rejection-prioritized checking order. Furthermore, exploiting server-oriented features, we propose a comprehensive GPU hardware design, augmented by a suite of GPU implementation optimizations to further amplify performance. Additionally, we present variants for sampling sparse polynomials, thereby streamlining our design. The deployment of our implementation on both server-grade and commercial GPUs demonstrates significant speedups, ranging from 284.9× to 485.3× against the CPU baseline, and an improvement of up to 60.9% compared to related work, affirming the effectiveness and efficiency of the proposed GPU architecture for ML-DSA signing procedure.
Shiyu Shen 0001, Hao Yang 0062, Yunlei Zhao
IEEE Trans. Dependable Secur. Comput.1
2024 Leveraging GPU in Homomorphic Encryption: Framework Design and Analysis of BFV Variants
abstract
Homomorphic Encryption (HE) enhances data security by enabling computations on encrypted data, advancing privacy-focused computations. The BFV scheme, a promising HE scheme, raises considerable performance challenges. Graphics Processing Units (GPUs), with considerable parallel processing abilities, offer an effective solution. In this work, we present an in-depth study on accelerating and comparing BFV variants on GPUs, including Bajard-Eynard-Hasan-Zucca (BEHZ), Halevi-Polyakov-Shoup (HPS), and recent variants. We introduce a universal framework for all variants, propose optimized BEHZ implementation, and first support HPS variants with large parameter sets on GPUs. We also optimize low-level arithmetic and high-level operations, minimizing instructions for modular operations, enhancing hardware utilization for base conversion, and implementing efficient reuse strategies and fusion methods to reduce computational and memory consumption. Leveraging our framework, we offer comprehensive comparative analyses. Performance evaluation shows a 31.9$\times$speedup over OpenFHE running on a multi-threaded CPU and 39.7% and 29.9% improvement for tensoring and relinearization over the state-of-the-art GPU BEHZ implementation. The leveled HPS variant records up to 4$\times$speedup over other variants, positioning it as a highly promising alternative for specific applications.
Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Lu Zhou 0002, Zhe Liu 0001, Yunlei Zhao
IEEE Trans. Computers1
2024 Phantom: A CUDA-Accelerated Word-Wise Homomorphic Encryption Library
abstract
Homomorphic encryption (HE) is a promising technique for privacy-preserving computations, especially the word-wise HE schemes that allow batching. However, the high computational overhead hinders the deployment of HE in real-word applications. GPUs are often used to accelerate execution, but a comprehensive performance comparison of different schemes on the same platform is still missing. In this work, we fill this gap by implementing three word-wise HE schemes BGV, BFV, and CKKS on GPU, with both theoretical and engineering optimizations. We enhance the hybrid key-switching technique, significantly reducing the computational and memory overhead. We explore several kernel fusing strategies to reuse data, resulting in reduced memory access and IO latency, and enhancing the overall performance. By comparing with the state-of-the-art works, we demonstrate the effectiveness of our implementation. Meanwhile, we introduce a unified framework that finely integrates our implementation of the three schemes, covering almost all scheme functions and homomorphic operations. We optimize the management of pre-computation, RNS bases, and memory in the framework, to provide efficient and low-latency data access and transfer. Based on this framework, we provide a thorough benchmark of the three schemes, which can serve as a reference for scheme selection and implementation in constructing privacy-preserving applications.
Hao Yang 0062, Shiyu Shen 0001, Wangchen Dai, Lu Zhou 0002, Zhe Liu 0001, Yunlei Zhao
IEEE Trans. Dependable Secur. Comput.2
2024 cuXCMP: CUDA-Accelerated Private Comparison Based on Homomorphic Encryption
abstract
Private comparison schemes constructed on homomorphic encryption offer the noninteractive and parallelizable features, and have advantages in communication bandwidth and performance. In this work, we propose cuXCMP, an extension of the privacy comparison scheme XCMP (AsiaCCS 2018). We address the relatively small input domain and the incompletely expressible output of XCMP, by modifying the encoding method and devising a constant term extraction (CTX) approach. Then, we describe a method for constructing privacy-preserving decision tree (PPDT) using this scheme. Considering the high computational overhead of CTX, we exploit the massive parallelism of the GPU to accelerate this function. Based on the results of the kernel profiling, we utilize several optimization techniques to improve the performance, including using multiple CUDA streams, reducing the grid dimension, kernel fusion, etc. By accelerating this function, we boost the execution time of the scheme and demonstrate 130× and 1.9× speedups for CTX and cuXCMP, respectively, as well as a 35% reduction in the evaluation time of PPDT.
Hao Yang 0062, Shiyu Shen 0001, Zhe Liu 0001, Yunlei Zhao
IEEE Trans. Inf. Forensics Secur.2
2024 High-Throughput GPU Implementation of Dilithium Post-Quantum Digital Signature
abstract
Digital signatures are fundamental building blocks in various protocols to provide integrity and authenticity. The development of the quantum computing has raised concerns about the security guarantees afforded by classical signature schemes. CRYSTALS-Dilithium is an efficient post-quantum digital signature scheme based on lattice cryptography and has been selected as the primary algorithm for standardization by the National Institute of Standards and Technology. In this work, we present a high-throughput GPU implementation of Dilithium. For individual operations, we employ a range of computational and memory optimizations to overcome sequential constraints, reduce memory usage and IO latency, address bank conflicts, and mitigate pipeline stalls. This results in high and balanced compute throughput and memory throughput for each operation. In terms of concurrent task processing, we leverage task-level batching to fully utilize parallelism and implement a memory pool mechanism for rapid memory access. We propose a dynamic task scheduling mechanism to improve multiprocessor occupancy and significantly reduce execution time. Furthermore, we apply asynchronous computing and launch multiple streams to hide data transfer latencies and maximize the computing capabilities of both CPU and GPU. Across all three security levels, our GPU implementation achieves over 160× speedups for signing and over 80× speedups for verification on both commercial and server-grade GPUs. This achieves microsecond-level amortized execution times for each task, offering a high-throughput and quantum-resistant solution suitable for a wide array of applications in real systems.
Shiyu Shen 0001, Hao Yang 0062, Wangchen Dai, Zhe Liu 0001, Yunlei Zhao
IEEE Trans. Parallel Distributed Syst.1
2023 Poster Abstract: CNN-guardian: Secure Neural Network Inference Acceleration on Edge GPU
abstract
The rapid development of AI applications powered by deep learning in edge devices boosts the opportunity for real-time health monitoring. To address the potential privacy concern in the inference phase, homomorphic encryption (HE) is an alternative solution that encrypts inference data without exposing raw data and has several distinct advantages, (i.e., single-round communication, lightweight bandwidth consumption, and non-interactive computation). However, the computational overhead on the current HE-based privacy-preserving inference necessitates a substantial amount of time, which is not feasible for some real-time applications on edge devices. To address this issue, we propose CNN-guardian, a unified and compact neural network structure for real-time inference in HE-based inference on edge GPU. CNN-guardian designs a HE-friendly neural network and GPU engine that optimizes HE operations to accelerate the inference in the HE domain.
Qipeng Xie, Hao Yang 0062, Linshan Jiang, Siyang Jiang, Shiyu Shen 0001, Salabat Khan, Zhe Liu 0001, Kaishun Wu
SenSys6
2023 CARM: CUDA-Accelerated RNS Multiplication in Word-Wise Homomorphic Encryption Schemes for Internet of Things
abstract
Homomorphic encryption (HE), which allows computation over encrypted data, has often been used to preserve privacy. However, the computationally heavy nature and complexity of network topologies make the deployment of HE schemes in the Internet of Things (IoT) scenario difficult. In this work, we propose CARM, the first optimized GPU implementation that covers BGV, BFV and CKKS, targeting for accelerating homomorphic multiplication using GPU in heterogeneous IoT systems. Our solution is suitable for accelerating RNS homomorphic multiplication on both high-performance and embedded GPUs, as it is a parametric and generic design and offers various trade-offs between resource and efficiency. We offer constant-time low-level arithmetic with minimum instructions and memory usage, as well as performance- and memory-prior configurations. Through this, we can provide more real-time evaluation results and relieve the computational pressure on cloud devices. We deploy our implementations on two GPUs. Compared to the CPU implementation, we achieve up to$378.4\times$,$234.5\times$, and$287.2\times$speedup for homomorphic multiplication of BGV, BFV, and CKKS on Tesla V100S, and$8.8\times$,$9.2\times$, and$10.3\times$on Jetson AGX Xavier, respectively.
Shiyu Shen 0001, Hao Yang 0062, Zhe Liu 0001, Yunlei Zhao
IEEE Trans. Computers1
2022 Parallel Small Polynomial Multiplication for Dilithium: A Faster Design and Implementation
abstract
The lattice-based signature scheme CRYSTALS-Dilithium is one of the two signature finalists in the third round NIST post-quantum cryptography (PQC) standardization project. For applications of low-power Internet-of-Things (IoT) devices, recent research efforts have been focusing on the performance optimization of PQC algorithms on embedded systems. In particular, performance optimization is more demanding for PQC signature algorithms that are usually significantly more time-consuming than PQC public-key encryption counterparts. For most cryptographic algorithms based on algebraic lattices including Dilithium, the fundamental and most time-consuming operation is polynomial multiplication over rings. For this computational task, number theoretic transform (NTT) is the most efficient multiplication method for NTT-friendly rings, and is now the typical technique for performing fast polynomial multiplications when implementing lattice-based PQC algorithms.
Jieyu Zheng, Shiyu Shen 0001, Chenxi Xue, Yunlei Zhao
ACSAC3
2022 Identity-based authenticated encryption with identity confidentiality
Shiyu Shen 0001, Yunlei Zhao
Theor. Comput. Sci.1
2022 Compact and Flexible KEM From Ideal Lattice
abstract
A remarkable breakthrough in mathematics in recent years is the proof of the long-standing conjecture: sphere packing in the$E_{8}$lattice is optimal in the sense of the best density for sphere packing in$\mathbb {R}^{8}$. In this work, we design a mechanism for asymmetric key consensus from noise (AKCN), referred to as AKCN-E8, for error correction and key consensus. As a direct application, we present a practical key encapsulation mechanism (KEM) from the ideal lattice based on the ring learning with errors (RLWE) problem. Compared with NewHope-KEM that was the second round candidate of the National Institute of Standards and Technology (NIST) post-quantum cryptography (PQC) standardization, our AKCN-E8 KEM scheme overcomes some limitations and shortcomings of NewHope-KEM. Compared with some other dominating KEM schemes based on the variants of LWE, specifically Kyber and Saber, AKCN-E8 has a comparable performance but enjoys much flexible shared-key sizes. Specifically, the key encapsulated by AKCN-E8-512 (resp., 768, 1024) has the size of 256 (resp., 384, 512) bits. Flexible key size renders us stronger security against quantum attacks, more powerful and economic ability of key transportation, and better matches the demand in interactive protocols like TLS where parties need to negotiate the security parameters including the shared key length.
Zhengzhong Jin, Shiyu Shen 0001, Yunlei Zhao
IEEE Trans. Inf. Theory2
2020 Number Theoretic Transform: Generalization, Optimization, Concrete Analysis and Applications
Zhichuang Liang, Shiyu Shen 0001, Yuantao Shi, Dongni Sun, Chongxuan Zhang, Guoyun Zhang, Yunlei Zhao, Zhixiang Zhao
Inscrypt2