Yuan Zhao 0015

dblp:65/2105-15 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
9since 2021 · last 2025
0009-0006-9633-9480ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 7 · 2 first-author · 3 since 2021Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ScaleOT: Privacy-utility-scalable Offsite-tuning with Dynamic LayerReplace and Selective Rank Compression
abstract
Offsite-tuning is a privacy-preserving method for tuning large language models (LLMs) by sharing a lossy compressed emulator from the LLM owners with data owners for downstream task tuning. This approach protects the privacy of both the model and data owners. However, current offsite tuning methods often suffer from adaptation degradation, high computational costs, and limited protection strength due to uniformly dropping LLM layers or relying on expensive knowledge distillation. To address these issues, we propose ScaleOT, a novel privacy-utility-scalable offsite-tuning framework that effectively balances privacy and utility. ScaleOT introduces a novel layerwise lossy compression algorithm that uses reinforcement learning to obtain the importance of each layer. It employs lightweight networks, termed harmonizers, to replace the raw LLM layers. By combining important original LLM layers and harmonizers in different ratios, ScaleOT generates emulators tailored for optimal performance with various model scales for enhanced privacy protection. Additionally, we present a rank reduction method to further compress the original LLM layers, significantly enhancing privacy with negligible impact on utility. Comprehensive experiments show that ScaleOT can achieve nearly lossless offsite tuning performance compared with full fine-tuning while obtaining better model privacy.
Zhaorui Tan, Tiandi Ye, Lichun Li, Yuan Zhao 0015, Wenyan Liu 0001, Wei Wang 0002, Jianke Zhu
AAAI5
2025 GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models
abstract
Kai Yao, Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao, Yixin Ji, Jianke Zhu, Wei Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao 0015, Yixin Ji, Jianke Zhu, Wei Wang 0002
ACL (1)7
2025 Gibbon: Faster Secure Two-party Training of Gradient Boosting Decision Tree
abstract
Gradient Boosting Decision Tree (GBDT) and its variants are widely used in industry. They have achieved remarkable success in numerous machine learning competitions and practical applications. Secure Multi-Party Computation (MPC) allows multiple data owners to compute a function jointly while keeping their input private. In this work, we present Gibbon, a secure two-party GBDT training framework on a vertically split dataset, where two data owners each hold different features of the same data samples. Compared with the state-of-the-art Squirrel (USENIX'Sec 2023), for most parameter settings, Gibbon achieves 2×-4× reduction in running time and 2×-3× reduction in communication.
Lichun Li, Zecheng Wu, Yuan Zhao 0015, Zhihao Li 0001
CCS3
2025 Privacy-Preserving Decision Graph Inference From Homomorphic Lookup Table
abstract
This paper studies MPC based decision graph inference (MDGI) where decision graphs (generalization of decision trees) are very popular machine learning models. In MDGI, a modeler holding a private model and a data owner holding private feature vectors jointly run model inference on each vector through an MPC protocol, which outputs the inference result without revealing the model or vector. Several noteworthy MDGI solutions have been proposed, but they remain unsatisfactory for large models due to the high communication cost of oblivious decision, the most complex component in MDGI. Oblivious decision securely evaluates binary tests (Boolean-valued functions) over features without revealing test type, parameter, feature index, or value. All constant-round oblivious decision protocols suffer from high communication costs due to bitwise encryption and transmission. Moreover, most support only one of the three common test types: threshold comparison, equality test, and set containment. We propose Homomorphic Lookup Table (HLT), a novel MPC technique for oblivious decision. HLT circumvents the bitwiseencryption and heavy-communication issue by adopting tablelookup-style computation and amortizing the cost of encrypted table transfer across many inferences. We carefully design the table structure and lookup rules, and integrate homomorphic encryption to optimize performance and support multiple test types. HLT achieves 78 × −4151× reduction in communication for oblivious decision, and is the first method to support all three common test types. Based on HLT, we design constant-round MDGI solutions for two widely used decision graphs: decision trees and scorecards. This is the first privacy-preserving solution for scorecards, and our decision tree solution reduces communication by 27 × −321×.
Lichun Li, Yuan Zhao 0015, Kai Bu, Linfeng Cheng
IEEE Trans. Dependable Secur. Comput.2
2025 GIF-FHE: A Comprehensive Implementation and Evaluation of GPU-Accelerated FHE With Integer and Floating-Point Computing Power
abstract
Fully Homomorphic Encryption (FHE) allows computations on encrypted data without revealing the plaintext, garnering significant interest from both academic and industrial communities. However, its broader adoption has been hindered by performance limitations. Consequently, researchers have turned to GPUs for efficient FHE implementation. Nevertheless, most have predominantly favored integer units due to their ease of use, overlooking the considerable computational potential of floating-point units in GPUs. Recognizing this untapped floating-point computational power, our paper introducesGIF-FHE, an extensive exploration and implementation of FHE, leveraging GPUs' integer and floating-point instructions for FHE acceleration. We develop a comprehensive suite of low-level and middle-level FHE primitives, offering multiple implementation variants with support for three word size configurations ($64/52/32$-bit). Particularly, we make innovative use of floating-point implementations, employing a novel methodology to efficiently leverage the floating-point unit's fused multiply-add (FMA) instructions. This represents the pioneering integration of floating-point units into FHE acceleration. To bridge our highly-optimized FHE primitives with practical applications, this paper also provides a high-level FHE implementation and interfaces that can be directly applied by upper-level applications such as neural network inference. Finally, we undertake a comprehensive experiment evaluation and comparison involving three types of arithmetic: FP64/INT64/INT32 with varying word size configurations and computation units. Notably, our fundamental function implementations consistently outperform counterparts on the same platform, achieving speedups ranging from$2.0\times$to$4.2\times$. In the context of CKKS FHE schemes, our homomorphic operation implementation surpasses the state-of-the-art GPU-based solution with a speedup of up to$3.8\times$, and exceeds the performance of the widely adopted CPU-based library, SEAL, with a remarkable speedup of over$300\times$.
Fangyu Zheng, Guang Fan 0001, Wenxu Tang, Yuan Zhao 0015, Jiankuo Dong, Jingqiang Lin 0001, Shoumeng Yan, Jiwu Jing
IEEE Trans. Parallel Distributed Syst.6
2023 Towards Faster Fully Homomorphic Encryption Implementation with Integer and Floating-point Computing Power of GPUs
abstract
Fully Homomorphic Encryption (FHE) allows computations on encrypted data without knowledge of the plaintext message and currently has been the focus of both academia and industry. However, the performance issue hinders its large-scale application, highlighting the urgent requirements of high-performance FHE implementations.With noticing the tremendous potential of GPUs in the field of cryptographic acceleration, this paper comprehensively investigates how to convert the available computing resources residing in GPUs into FHE workhorses, and implement a full set of low-level and middle-level FHE primitives based on two arithmetic units (i.e., INT32 and FP64 units) with three types of data precision (i.e., INT32, INT64 and FP64). This paper gives a comprehensive evaluation and comparison based on each road-map. Our implementations of fundamental functions outperform the implementations on the same platform by 1.7× to 16.7×. Taking CKKS FHE schemes as a case study, our implementation of homomorphic multiplication achieves 3.2× speedup over the state-of-the-art GPU-based implementation, even considering the difference of platforms. The detailed evaluation and comparison of this paper would offer a vital reference for the follow-up work to choose appropriate underlying arithmetic units and important primitive optimizations in GPU-based FHE implementations.
Guang Fan 0001, Fangyu Zheng, Lipeng Wan 0002, Yuan Zhao 0015, Jiankuo Dong, Yuewu Wang, Jingqiang Lin 0001
IPDPS5
2023 Topgun: An ECC Accelerator for Private Set Intersection
abstract
Elliptic Curve Cryptography (ECC), one of the most widely used asymmetric cryptographic algorithms, has been deployed in Transport Layer Security (TLS) protocol, blockchain, secure multiparty computation, and so on. As one of the most secure ECC curves, Curve25519 is employed by some secure protocols, such as TLS 1.3 and Diffie-Hellman Private Set Intersection (DH-PSI) protocol. High-performance implementation of ECC is required, especially for the DH-PSI protocol used in privacy-preserving platform. Point multiplication, the chief cryptographic primitive in ECC, is computationally expensive. To improve the performance of DH-PSI protocol, we propose Topgun, a novel and high-performance hardware architecture for point multiplication over Curve25519. The proposed architecture features a pipelined Finite-field Arithmetic Unit and a simple and highly efficient instruction set architecture. Compared to the best existing work on Xilinx Zynq 7000 series FPGA, our implementation with one Processing Element can achieve 3.14× speedup on the same device. To the best of our knowledge, our implementation appears to be the fastest among the state-of-the-art works. We also have implemented our architecture consisting of 4 Compute Groups, each with 16 PEs, on an Intel Agilex AGF027 FPGA. The measured performance of 4.48 Mops/s is achieved at the cost of 86 Watts power, which is the record-setting performance for point multiplication over Curve25519 on FPGAs.
Guiming Wu, Qianwen He, Jiali Jiang, Zhenxiang Zhang, Yuan Zhao 0015, Yinchao Zou, Jie Zhang 0144, Changzheng Wei, Ying Yan 0002, Hui Zhang 0002
ACM Trans. Reconfigurable Technol. Syst.5
2022 A High-Performance Hardware Architecture for ECC Point Multiplication over Curve25519
abstract
As one of the most secure ECC curves, Curve25519 is employed by some secure protocols, such as TLS 1.3, IRTF’s RFC7748, Diffie-Hellman Private Set Intersection (DH-PSI) protocol, etc. High performance implementation of ECC is required, especially for the DH-PSI protocol. Point multiplication, the chief cryptographic primitive in ECC, is computationally expensive. To improve the performance of DH-PSI protocol, we propose a novel and high-performance hardware architecture for point multiplication over Curve25519. The proposed architecture features a pipelined Finite-field Arithmetic Unit (FAU) and a simple and highly efficient instruction set architecture (ISA). Compared to the best existing work on Xilinx Zynq 7000 series FPGA, our implementation with one Processing Element (PE) can achieve 3.14x speedup on the same device. To the best of our knowledge, our implementation appears to be the fastest among the state-of-the-art works. We also have implemented our proposed architecture consisting of 4 Compute Groups (CGs), each with 16 PEs, on an Intel Agilex AGF027 FPGA. The experimental results show the peak performance of 4.52 Mops/s (million point multiplication operations per seconds) can be achieved. Moreover, the measured performance of 4.48 Mops/s is achieved, with the PE utilization of 99% and at the cost of 86 Watts power, which is the record-setting performance for point multiplication over Curve25519 on FPGAs.
Guiming Wu, Qianwen He, Jiali Jiang, Zhenxiang Zhang, Yuan Zhao 0015, Yinchao Zou
FCCM6
2021 VIRSA: Vectorized In-Register RSA Computation with Memory Disclosure Resistance
Yu Fu 0007, Wei Wang 0314, Lingjia Meng, Qiongxiao Wang, Yuan Zhao 0015, Jingqiang Lin 0001
ICICS (1)5
2017 Utilizing the Double-Precision Floating-Point Computing Power of GPUs for RSA Acceleration
abstract
Asymmetric cryptographic algorithm (e.g., RSA and Elliptic Curve Cryptography) implementations on Graphics Processing Units (GPUs) have been researched for over a decade. The basic idea of most previous contributions is exploiting the highly parallel GPU architecture and porting the integer-based algorithms from general-purpose CPUs to GPUs, to offer high performance. However, the great potential cryptographic computing power of GPUs, especially by the more powerful floating-point instructions, has not been comprehensively investigated in fact. In this paper, we fully exploit the floating-point computing power of GPUs, by various designs, including the floating-point-based Montgomery multiplication/exponentiation algorithm and Chinese Remainder Theorem (CRT) implementation in GPU. And for practical usage of the proposed algorithm, a new method is performed to convert the input/output between octet strings and floating-point numbers, fully utilizing GPUs and further promoting the overall performance by about 5%. The performance of RSA-2048/3072/4096 decryption on NVIDIA GeForce GTX TITAN reaches 42,211/12,151/5,790 operations per second, respectively, which achieves 13 times the performance of the previous fastest floating-point-based implementation (published in Eurocrypt 2009). The RSA-4096 decryption precedes the existing fastest integer-based result by 23%.
Jiankuo Dong, Fangyu Zheng, Wuqiong Pan, Jingqiang Lin 0001, Jiwu Jing, Yuan Zhao 0015
Secur. Commun. Networks6
2016 PhiRSA: Exploiting the Computing Power of Vector Instructions on Intel Xeon Phi for RSA
Yuan Zhao 0015, Wuqiong Pan, Jingqiang Lin 0001, Peng Liu 0005, Fangyu Zheng
SAC1
2016 RegRSA: Using Registers as Buffers to Resist Memory Disclosure Attacks
Yuan Zhao 0015, Jingqiang Lin 0001, Wuqiong Pan, Fangyu Zheng, Ziqiang Ma
SEC1
2014 Exploiting the Floating-Point Computing Power of GPUs for RSA
Fangyu Zheng, Wuqiong Pan, Jingqiang Lin 0001, Jiwu Jing, Yuan Zhao 0015
ISC5