Yu Fu 0007

dblp:09/3263-7 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-9895-3816ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 9 · 4 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GRASP: Accelerating Hash-Based PQC Performance on GPU Parallel Architecture
abstract
SPHINCS+, one of the Post-Quantum Cryptography Digital Signature Algorithms (PQC-DSA) selected by NIST in the third round, features very short public and private key lengths but faces significant performance challenges compared to other post-quantum cryptographic schemes, limiting its suitability for real-world applications. In scenarios involving a large number of concurrent signing or verification tasks, these performance bottlenecks become particularly critical. To address these challenges, we propose the GPU-based paRallel Accelerated SPHINCS+(GRASP), which leverages GPU technology to enhance the efficiency of SPHINCS+signing and verification processes. We propose an adaptable parallelization strategy for SPHINCS+, analyzing its signing and verification processes to identify critical sections for efficient parallel execution. Utilizing CUDA, we perform bottom-up optimizations, focusing on memory access patterns and hypertree computation, to enhance GPU resource utilization. These efforts, combined with kernel fusion technology, result in significant improvements in throughput and overall performance. Compared to previous works, our approach achieves the highest occupancy. Extensive experimentation demonstrates that our optimized CUDA implementation of SPHINCS+achieves superior performance. Specifically, our GRASP scheme delivers throughput improvements ranging from 1.09× to 3.45× compared to state-of-the-art GPU-based solutions and surpasses the NIST reference implementation by over three orders of magnitude, highlighting a significant performance advantage.
Yijing Ning, Jiankuo Dong, Jingqiang Lin 0001, Fangyu Zheng, Yu Fu 0007, Fu Xiao 0001
IEEE Trans. Computers5
2026 TFMD: General and Fast Secure Neural Network Inference Framework With Threshold FHE
abstract
Secure neural network inference is the privacy-preserving inference method that protects the model parameters and user’s private input. Previous works have constructed two-party, three-party and four-party secure inference schemes. However, these schemes allow only one corrupted party. Also, the interaction protocol between different parties is customized based on the number of participants. If the number of participants increases or decreases, the protocol needs to be redesigned. Another problem is that current protocols for non-linear functions still have large computation overhead. In this work, we present TFMD, a general and fast secure neural network inference framework with semi-honest security. TFMD is built based on threshold fully homomorphic encryption (FHE), and is suitable for the outsourced computation scenario. Concretely, TFMD designs general secure computation protocols for non-linear functions. Our protocols support arbitrarynparticipants, and allow at mostn– 1 corrupted parties. Further, TFMD constructs a novel secure neural network inference framework. TFMD employs FHE with computation-friendly coefficient encoding to quickly calculate linear functions, and employs our proposed protocol to calculate ReLU. Experiments illustrate that TFMD is both efficient and scalable. Even in the three-party setting, the online phase of our inference is 2.1× faster than CrypTFlow (S&P’20).
Yu Fu 0007, Yijing Ning, Jingqiang Lin 0001, Dengguo Feng
IEEE Trans. Inf. Forensics Secur.1
2025 Octopus: Fast Homomorphic Convolution for Secure Neural Network Inference
abstract
Secure two-party neural network (2PC-NN) inference is a privacy-preserving inference method that protects the client's input and the server's model parameters. While addressing privacy concerns, it also incurs considerable over-heads. In this work, we propose Octopus, a faster and more communication-efficient 2PC-NN system than prior works. Octopus designs an optimized encoding method for fast homomorphic convolution, and further constructs homomorphic encryption-based convolutional computation protocol. Compared with the original coefficient encoding proposed by Cheetah, our method significantly reduces the resulting ciphertexts through packing output channels, thereby saving the communication cost and end-to-end execution time. Moreover, Octopus proposes an encoding-motivated fine tuning technique for convolutional neural networks, which fully utilizes the feature of coefficient encoding to adaptively adjust the neural network structure to maximize performance with negligible accuracy loss. We apply Octopus to the widely used model ResNet on CIFAR-10 and ImageNet dataset. Experiments illustrate that Octopus has obvious improvement compared with the state-of-the-art approaches, achieving a speedup of up to 2.75×, and reduces communication overhead by up to 7.19× for convolutions. As for secure inference, compared with Cheetah (resp., CrypTFlow2), Octopus demonstrates 1.41× (resp., 13.20×) lower communication cost and 1.25× (resp., 7.03×) faster execution time under a WAN setting.
Yu Fu 0007, Tianshi Xu, Cheng Hong 0001, Meng Li 0004, Wei Wang 0314, Dengguo Feng, Jingqiang Lin 0001
ACSAC2
2025 Exploring the Root Store Usage in TLS-Based Applications
Yuxiang Shen, Wei Wang 0314, Shushang Wen, Yu Fu 0007, Yunhao Jia, Jingqiang Lin 0001
Inscrypt (2)4
2025 Swift: Fast Secure Neural Network Inference With Fully Homomorphic Encryption
abstract
With the widespread use of machine learning (ML), privacy concerns during neural network inference are attracting growing attention. Secure two-party neural network (2PC-NN) inference is the privacy-preserving inference method, which allows client to obtain the inference result without disclosing client’s input to the server. The server’s model parameters are also confidential to the client. However, current 2PC-NN inference schemes still have large overhead, especially for non-linear functions. In this paper, we present Swift, a fast secure 2PC-NN inference scheme based on fully homomorphic encryption (FHE) and secret sharing (SS). FHE protects the input and model parameters in linear functions, while SS is integrated to protect the non-linear functions. Concretely, Swift integrates FHE and SS to design secure and efficient non-linear protocols used for ReLU and max pooling. To further optimize performance, Swift employs FHE with computation-friendly coefficient encoding for fast execution of linear functions, and SIMD encoding for non-linear functions. Swift constructs efficient encoding conversion protocol between the coefficient-encoded ciphertext and the SIMD-encoded ciphertext. Finally, Swift achieves secure neural network inference framework for MNIST dataset. Compared with Cheetah (USENIX 2022), the execution time of ReLU, max pooling, secure inference under a WAN setting improves$7.4\times $,$13.3\times $,$1.9\times $, respectively.
Yu Fu 0007, Yijing Ning, Tianshi Xu, Meng Li 0004, Jingqiang Lin 0001, Dengguo Feng
IEEE Trans. Inf. Forensics Secur.1
2025 HTM-PQC: Hardening Cryptography Keys Under the Trend of Post-Quantum Cryptography Migration on Industrial Internet
abstract
With the rapid expansion of Industry 4.0 technology, the proliferation of large-scale devices faces increasingly severe cyber threats, underscoring the critical importance of cryptographic technology for secure communication and authentication. However, cryptographic systems, as the bedrock of security, have faced a barrage of attacks in recent years, including potential threats from quantum computing and memory disclosure vulnerabilities. In this article, we focus on enhancing the security of two standard quantum-safe cryptographic algorithms, Dilithium and eXtended Merkle signature scheme (XMSS), by leveraging hardware transactional memory (HTM) to create a secure operational environment. Unlike traditional cryptography such as Rivest–Shamir–Adleman (RSA) and elliptic curve cryptography (ECC), Dilithium, and XMSS involve more and larger sensitive variables, rendering conventional solutions inadequate. By conducting a comprehensive sensitivity analysis of variables within the abovementioned algorithms, we confine sensitive operations to transactional execution regions and employ transaction-splitting technology for efficiency. Our prototype, utilizing Intel transactional synchronization extension (TSX), demonstrates robust protection against memory disclosure attacks with acceptable performance overheads. Notably, our security-enhanced Dilithium and XMSS software implementations, recommended by NIST, achieve an average throughput factor of 0.75 compared to the (unprotected) reference implementations.
Lingjia Meng, Yu Fu 0007, Fangyu Zheng, Ziqiang Ma, Jiankuo Dong, Jingqiang Lin 0001
IEEE Trans. Ind. Informatics2
2023 Protecting Private Keys of Dilithium Using Hardware Transactional Memory
Lingjia Meng, Yu Fu 0007, Fangyu Zheng, Ziqiang Ma, Dingfeng Ye, Jingqiang Lin 0001
ISC2
2023 RegKey: A Register-based Implementation of ECC Signature Algorithms Against One-shot Memory Disclosure
abstract
To ensure the security of cryptographic algorithm implementations, several cryptographic key protection schemes have been proposed to prevent various memory disclosure attacks. Among them, the register-based solutions do not rely on special hardware features and offer better applicability. However, due to the size limitation of register resources, the performance of register-based solutions is much worse than conventional cryptosystem implementations without security enhancements. This paper presents RegKey, an efficient register-based implementation of ECC (elliptic curve cryptography) signature algorithms. Different from other schemes that protect the whole cryptographic operations, RegKey only uses CPU registers to execute simple but critical operations, significantly reducing the usage of register resources and performance overheads. To achieve this goal, RegKey splits the ECC signing into two parts, (1) complex elliptic curve group operations on non-sensitive data in main memory as normal implementations, and (2) simple prime field operations on sensitive data inside CPU registers. RegKey guarantees the plaintext private key and random number used for signing only appear in registers to effectively resist one-shot memory disclosure attacks such as cold-boot attacks and warm-boot attacks, which are usually launched by physically accessing the victim machine to acquire partial or even entire memory data but only once. Compared with existing cryptographic key protection schemes, the performance of RegKey is greatly improved. Regkey is applicable to different platforms because it does not rely on special CPU hardware features. Since RegKey focuses on one-shot memory disclosure instead of persistent software-based attacks, it works as a choice suitable for embedded devices or offline machines where physical attacks are the main threat.
Yu Fu 0007, Jingqiang Lin 0001, Dengguo Feng, Wei Wang 0314
ACM Trans. Embed. Comput. Syst.1
2021 SMCOS: Fast and Parallel Modular Multiplication on ARM NEON Architecture for ECC
Wei Wang 0314, Jingqiang Lin 0001, Yu Fu 0007, Lingjia Meng, Qiongxiao Wang
Inscrypt4
2021 VIRSA: Vectorized In-Register RSA Computation with Memory Disclosure Resistance
Yu Fu 0007, Wei Wang 0314, Lingjia Meng, Qiongxiao Wang, Yuan Zhao 0015, Jingqiang Lin 0001
ICICS (1)1
2020 Improving the Effectiveness of Grey-box Fuzzing By Extracting Program Information
abstract
Fuzzing has been widely adopted as an effective techniques to detect vulnerabilities in softwares. However, existing fuzzers suffer from the problems of generating excessive test inputs that either cannot pass input validation or are ineffective in exploring unvisited regions in the program under test (PUT). To tackle these problems, we propose a greybox fuzzer called MuFuzzer based on AFL, which incorporates two heuristics that optimize seed selection and automatically extract input formatting information from the PUT to increase the chance of generating valid test inputs, respectively. In particular, the first heuristic collects the branch coverage and execution information during a fuzz session, and utilizes such information to guide fuzzing tools in selecting seeds that are fast to execute, small in size, and more importantly, more likely to explore new behaviors of the PUT for subsequent fuzzing activities. The second heuristic automatically identifies string comparison operations that the PUT uses for input validation, and establishes a dictionary with string constants from these operations to help fuzzers generate test inputs that have higher chances to pass input validation. We have evaluated the performance of MuFuzzer, in terms of code coverage and bug detection, using a set of realistic programs and the LAVA-M test bench. Experiment results demonstrate that MuFuzzer is able to achieve higher code coverage and better or comparative bug detection performance than state-of-the-art fuzzers.
Yu Fu 0007, Siming Tong, Liang Cheng 0004, Yang Zhang 0021, Dengguo Feng
TrustCom1
2015 Improving Accuracy of Static Integer Overflow Detection in Binary
Yang Zhang 0021, Xiaoshan Sun, Yi Deng 0002, Liang Cheng 0004, Shuke Zeng, Yu Fu 0007, Dengguo Feng
RAID6