Yijing Ning

dblp:382/0925 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0006-6534-4260ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 GRASP: Accelerating Hash-Based PQC Performance on GPU Parallel Architecture
abstract
SPHINCS+, one of the Post-Quantum Cryptography Digital Signature Algorithms (PQC-DSA) selected by NIST in the third round, features very short public and private key lengths but faces significant performance challenges compared to other post-quantum cryptographic schemes, limiting its suitability for real-world applications. In scenarios involving a large number of concurrent signing or verification tasks, these performance bottlenecks become particularly critical. To address these challenges, we propose the GPU-based paRallel Accelerated SPHINCS+(GRASP), which leverages GPU technology to enhance the efficiency of SPHINCS+signing and verification processes. We propose an adaptable parallelization strategy for SPHINCS+, analyzing its signing and verification processes to identify critical sections for efficient parallel execution. Utilizing CUDA, we perform bottom-up optimizations, focusing on memory access patterns and hypertree computation, to enhance GPU resource utilization. These efforts, combined with kernel fusion technology, result in significant improvements in throughput and overall performance. Compared to previous works, our approach achieves the highest occupancy. Extensive experimentation demonstrates that our optimized CUDA implementation of SPHINCS+achieves superior performance. Specifically, our GRASP scheme delivers throughput improvements ranging from 1.09× to 3.45× compared to state-of-the-art GPU-based solutions and surpasses the NIST reference implementation by over three orders of magnitude, highlighting a significant performance advantage.
Yijing Ning, Jiankuo Dong, Jingqiang Lin 0001, Fangyu Zheng, Yu Fu 0007, Fu Xiao 0001
IEEE Trans. Computers1
2026 TFMD: General and Fast Secure Neural Network Inference Framework With Threshold FHE
abstract
Secure neural network inference is the privacy-preserving inference method that protects the model parameters and user’s private input. Previous works have constructed two-party, three-party and four-party secure inference schemes. However, these schemes allow only one corrupted party. Also, the interaction protocol between different parties is customized based on the number of participants. If the number of participants increases or decreases, the protocol needs to be redesigned. Another problem is that current protocols for non-linear functions still have large computation overhead. In this work, we present TFMD, a general and fast secure neural network inference framework with semi-honest security. TFMD is built based on threshold fully homomorphic encryption (FHE), and is suitable for the outsourced computation scenario. Concretely, TFMD designs general secure computation protocols for non-linear functions. Our protocols support arbitrarynparticipants, and allow at mostn– 1 corrupted parties. Further, TFMD constructs a novel secure neural network inference framework. TFMD employs FHE with computation-friendly coefficient encoding to quickly calculate linear functions, and employs our proposed protocol to calculate ReLU. Experiments illustrate that TFMD is both efficient and scalable. Even in the three-party setting, the online phase of our inference is 2.1× faster than CrypTFlow (S&P’20).
Yu Fu 0007, Yijing Ning, Jingqiang Lin 0001, Dengguo Feng
IEEE Trans. Inf. Forensics Secur.3
2026 X2O: Cross Parallel Optimization of the CROSS Post-Quantum Scheme on GPU
abstract
The CROSS Digital Signature Algorithm (DSA), currently a second-round candidate in the NIST standardization process for additional post-quantum digital signatures, offers compact public keys and strong security guarantees rooted in the code-based Restricted Syndrome Decoding Problem (R-SDP) and its variant R-SDP(G). Despite its strong theoretical foundation and practical significance, existing CPU-based implementations of CROSS exhibit evident performance limitations, while its potential for high-throughput acceleration on GPU architectures remains insufficiently investigated. In this work, we present X2O, the first systematically optimized GPU implementation framework for CROSS on NVIDIA GPUs. X2O introduces a novel cross-parallel architecture that integrates both horizontal and vertical parallelism to fully exploit the massive concurrency of modern GPU platforms. The framework incorporates a series of targeted optimizations, including fine-grained thread scheduling, optimized memory access patterns, hash function tuning, and GPU-efficient tree construction. Experimental results on a NVIDIA RTX 4090 demonstrate the efficiency of our design, achieving up to 1,082,904 signature generations and 1,589,595 verifications per second at NIST security level 1. Compared to the official AVX-optimized CPU implementation, our GPU-based approach achieves up to 120× speedup, establishing a new performance benchmark for CROSS and demonstrating the viability of high-throughput, post-quantum digital signatures on parallel computing platforms.
Yijing Ning, Jiankuo Dong, Jingqiang Lin 0001, Fu Xiao 0001
IEEE Trans. Inf. Forensics Secur.1
2025 ML-Cube: Accelerating Module-Lattice-Based Cryptography using Machine Learning Accelerators with a Memory-Less Design
abstract
The rapid advancement of AI technologies has led to a dramatic surge in computational demands, driving significant breakthroughs in ML accelerators. The powerful performance of these accelerators has attracted the attention of cryptography researchers, and recent studies have begun to explore their use in accelerating cryptographic operations. However, treating these accelerators as black boxes leads to high latency, and strict concurrency requirements, which hinder their practical deployment. In this paper, we go beyond the black-box treatment of ML accelerators and introduce ML-Cube (ML3), a novel memory-less framework that leverages ML accelerators to implement module-lattice-based PQC, FIPS 203 ML-KEM, and FIPS 204 ML-DSA. The performance benefits of ML-Cube arise from our thorough analysis of ML accelerator internals. Rather than treating the accelerators as black boxes, we dissect their operating mechanisms and design tailored mathematical transformations for cryptographic acceleration. This enables memory-less (I)NTT and polynomial multiplication that minimizes external memory dependencies and reduces latency. We further address the high latency and excessive parallelism demands of traditional SIMT-based implementations by fully parallelizing both ML-KEM and ML-DSA schemes. Our experiments show that our Tensor Core-based (I)NTT achieves a 2.03x--3.56x speedup over a highly-optimized CUDA-core implementation. Moreover, our memory-less polynomial multiplication attains a 10x speedup, and the full ML-KEM reaches up to a 3.58x speedup with only less than one-tenth of the latency compared with SOTA approach (CHES '24). Additionally, our enhanced ML-DSA implementation offers a 30% to 55% throughput improvement over the previous SOTA methods (TDSC '24) under the server-oriented model. Importantly, by confining core computations within registers, our approach inherently mitigates memory disclosure and cache-based side-channel attacks, thereby enhancing overall security.
Fangyu Zheng, Zhuoyu Xie, Wenxu Tang, Guang Fan 0001, Yijing Ning, Yi Bian 0001, Jingqiang Lin 0001, Jiwu Jing
CCS6
2025 Swift: Fast Secure Neural Network Inference With Fully Homomorphic Encryption
abstract
With the widespread use of machine learning (ML), privacy concerns during neural network inference are attracting growing attention. Secure two-party neural network (2PC-NN) inference is the privacy-preserving inference method, which allows client to obtain the inference result without disclosing client’s input to the server. The server’s model parameters are also confidential to the client. However, current 2PC-NN inference schemes still have large overhead, especially for non-linear functions. In this paper, we present Swift, a fast secure 2PC-NN inference scheme based on fully homomorphic encryption (FHE) and secret sharing (SS). FHE protects the input and model parameters in linear functions, while SS is integrated to protect the non-linear functions. Concretely, Swift integrates FHE and SS to design secure and efficient non-linear protocols used for ReLU and max pooling. To further optimize performance, Swift employs FHE with computation-friendly coefficient encoding for fast execution of linear functions, and SIMD encoding for non-linear functions. Swift constructs efficient encoding conversion protocol between the coefficient-encoded ciphertext and the SIMD-encoded ciphertext. Finally, Swift achieves secure neural network inference framework for MNIST dataset. Compared with Cheetah (USENIX 2022), the execution time of ReLU, max pooling, secure inference under a WAN setting improves$7.4\times $,$13.3\times $,$1.9\times $, respectively.
Yu Fu 0007, Yijing Ning, Tianshi Xu, Meng Li 0004, Jingqiang Lin 0001, Dengguo Feng
IEEE Trans. Inf. Forensics Secur.3