Yinghao Yang 0001

dblp:318/7098-1 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-0551-5703ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-author · 11 since 2021
YearPublicationVenuePosition
2026 LayerTEE: Decoupled Memory Protection for Scalable Multilayer Communication on RISC-V
abstract
The Trusted Execution Environment (TEE) has been widely implemented by modern hardware vendors to protect security and privacy-sensitive applications and data, such as Intel SGX/TDX, ARM TrustZone, AMD SEV, and RISC-V Penglai. However, existing TEE systems face challenges in balancing memory isolation among security, performance, and scalability requirements. Segment-based memory isolation mechanisms, like RISC-V PMP, struggle to scale effectively to the large number of segments needed for confidential cloud and data center environments. On the other hand, table-based isolation methods, such as page tables, combine address translation with memory protection, leading to inefficient cross-enclave communication and potential security vulnerabilities like Rowhammer attacks. This paper introduces a novel TEE system, which decouples memory protection (to segments) from address translation (to page tables). This design improves communication performance by dynamically adjusting memory protection capabilities, without sacrificing application compatibility. LayerTEE enhances enclave security and scalability by designing a multi-layer segment-based isolation mechanism. We have built a prototype of based on FPGA, incorporating hardware extensions and software support. The evaluation demonstrates that significantly surpasses existing TEE solutions, achieving three orders of magnitude lower communication latency and 10x greater scalability while maintaining robust security guarantees.
Shangjie Pan, Yinghao Yang 0001, Xuanyao Peng, Xiquan Zhao, Dong Du 0003, Yubin Xia, Xiaowei Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Hypnos: Memory Efficient Homomorphic Processing Unit
abstract
Fully Homomorphic Encryption (FHE) introduces a novel paradigm in privacy-preserving computation, extending its applicability to various scenarios. However, Operating on encrypted data imposes significant computational challenges, particularly elevating data transmission and memory access demands. Consequently, developing an efficient system storage architecture becomes vital for FHE-specific architectures. Traditional FHE accelerators use a Host+ACC topology, often focusing on enhancing computational performance and efficient using of on-chip caches, with the assumption that very large volumes of encrypted data are already present in the accelerator’s memory while neglecting the inefficiencies of the unavoidable PCIe bus. In this paper, we propose Hypnos—a memory-efficient homomorphic encryption processing unit. In Hypnos, we abstract operators from FHE schemes into commands suitable for memory-efficient processing units and combine them with a homomorphic encryption paged memory management system designed for memory access, significantly reducing the memory access and execution time of homomorphic encryption applications. We implement Hypnos on the QianKun FPGA Card and highlight the following results: (1) outperforms SOTA ASIC and FPGA solutions by 2.58× and 4.43× (2) the communication overhead is reduced by 3.78 × compared to traditional architectures; (3) up to 27.6× and 19.06× energy efficiency improvement compared to ASIC-based CraterLake and FPGA-based Poseidon for ResNet-20 respectively.
Yinghao Yang 0001, Xiaowei Li 0001
DAC2
2025 Ares: High Performance Near-Storage Accelerator for FHE-based Private Set Intersection
abstract
Nowadays, the importance of data privacy protection has grown significantly. Privacy Set Intersection (PSI) based on Fully Homomorphic Encryption (FHE) is widely applied in various privacy protection scenarios, such as federated learning and password verification. Nevertheless, the substantial computational demands of FHE and the vast scale of databases in PSI result in inefficient processing, thereby necessitating specialized accelerator architectures to enhance usability. Current general-purpose FHE accelerators do not adequately address the unique requirements of PSI applications, leading to suboptimal data handling and underutilization of hardware, which impedes their effective deployment for PSI acceleration. This paper introduces Ares, a practical hardware-software co-designed FHE-based PSI FPGA accelerator. We propose Lazy Relinearization to optimize redundant calculations in PSI and reduce computational complexity without changing the PSI protocol. At the same time, through the analysis and decoupling of the PSI computing pattern, we design an efficient hardware acceleration architecture that fully utilizes the bandwidth and computing resources of the hardware to achieve excellent acceleration performance. We highlight the following result: (1) a $47.99 \times$ speedup relative to CPU; (2) performance improvements of $1.79 \times$ and $1.93 \times$ over the state-of-theart FPGA FHE accelerators, Poseidon and FAB, respectively; (3) achieves $7.96 \times$ and $10.95 \times$ energy efficiency improvement compared to Poseidon and FAB, respectively.
Yinghao Yang 0001, Jinkai Zhang, Xiaowei Li 0001
DAC2
2025 Hydra: Scale-out FHE Accelerator Architecture for Secure Deep Learning on FPGA
abstract
Deep learning, including Convolutional Neural Network (CNN) and Large Language Model (LLM), under Fully Homomorphic Encryption (FHE) is very computationally intensive because of the burdensome computations like ciphertext convolution and matrix multiplication, non-linear layers, and bootstrapping. Existing FHE accelerators focus on the high throughput computational units, stacking parallelized clusters to maximize ciphertext inference performance. Nevertheless, this design philosophy cannot leverage the substantial parallelism at the application level and is not scalable for further performance enhancement by simply adding additional compute nodes to cope with the ever-increasing model sizes in the future. In this paper, we propose the high-performance FHE acceleration architecture in a “scale-out” manner for secure deep learning, termed as Hydra. It supports the multi-server scaling and arbitrary computational nodes theoretically, each handling a portion of the deep learning model governed by the central scheduling mechanism on the host server. Hydra exhibits excellent scalability and delivers outstanding performance across a range of compute resource sizes. We highlight the following results: (1) up to $74 \times$ and $160 \times$ speedup over the SOTA single card accelerator Poseidon and FAB; (2) outperforms 8-card FAB-2 by $12 \times$ to $21 \times$ for FHE-based CNNs and LLMs; (3) outperforms SOTA ASIC accelerators, CraterLake and SHARP, by $8.1 \times$ and $2.5 \times$ for LLM OPT-6.7B, and achieves comparable or superior energy efficiency under the same chip technology.
Yinghao Yang 0001, Xicheng Xu, Xiaowei Li 0001
HPCA1
2025 RTPU: Unifying Non-Private and Private Inference with Reconfigurable Architecture
abstract
With the rise of fully homomorphic encryption-based private inference, data centers are anticipated to simultaneously handle two disparate computational demands: plaintext-based non-private inference (NPI) and ciphertext-based private inference (PI). Unfortunately, current solutions face challenges in addressing this trend. They either depend on costly, inflexible dedicated accelerators or utilize general-purpose hardware with inferior performance. This limitation underscores the urgent need for a unified architecture capable of serving both normal and privacy-sensitive users with high efficiency.However, the fundamental disparities in computation patterns and resource management between NPI and PI make their architectural fusion intricate. To bridge this gap, we explore their inherent similarities and apply fine-grained reconfiguration to maximize resource sharing. We propose RTPU, a reconfigurable multi-core architecture that can seamlessly switch between tensor-based plaintext and polynomial ring-based ciphertext computations. Building upon its reconfigurable computing fabric and parallelization mechanism, we introduce a kernel group-based scheduling strategy to optimize hardware utilization and QoS. Experimental results show that: i) The RTPU architecture achieves near-ASIC performance and beyond-ASIC flexibility with substantial silicon reuse between NPI and PI. ii) The RTPU scheduler sustains high resource utilization for multi-tenant workloads with varying privacy requirements.
Fuping Li, Ying Wang 0001, Yinghao Yang 0001, Yibo Du, Huawei Li 0001, Yinhe Han 0001, Xiaowei Li 0001
ICCAD3
2025 Uranus: Ultra-efficient Acceleration Architecture for the Privacy Inference of Graph Neural Networks
abstract
Graph Neural Networks (GNNs) are increasingly applied across various domains, including social media and recommendation systems. However, data privacy protection has become a critical issue for GNN applications. Fully Homomorphic Encryption (FHE), which enables computation on encrypted data, has emerged as a mainstream solution for secure GNN inference. However, due to the high computational costs of FHE-based GNNs, dedicated accelerators are necessary to make them practical. Existing CKKS-based GNN algorithms require large encryption parameters, imposing high demands on hardware resources. The state-of-the-art design PPGNN is hardware friendly on memory by combining CKKS with TFHE, but TFHE’s SISD computing characteristic causes a dramatic increase in computational overhead, limiting its performance benefits. This paper proposes a novel software-hardware code-signed GNN accelerator architecture, Uranus, which integrates a hardware-friendly algorithmic framework with a matching accelerator design. The inference framework leverages CKKS and BFV schemes to enable efficient linear layer computations and SIMD-based nonlinear calculations with reduced encryption parameters suitable for hardware constraints. Additionally, the highly reconfigurable hardware accelerator maximizes utilization and performance. Key results demonstrate the following: (1) up to 92× speedup compared to CKKS accelerators; (2) achieves 3.3× to 4.2× performance improvement over the SOTA PPGNN architecture; (3) over 119× in energy efficiency improvement compared to PPGNN.
Xicheng Xu, Yinghao Yang 0001, Fuyao Liu, Xiaowei Li 0001
ICCAD2
2025 SecNPU: Securing LLM Inference on NPU
abstract
In the era of prevalent large language models (LLMs), efficient LLM inference systems deployed on the neural processing units (NPUs) have gained widespread adoption. During NPUbased LLM inference, both user privacy inputs and proprietary model parameters require stringent protection. While traditional trusted execution environments (TEEs) can be applied to NPU inference processes, we identify that they introduce challenging security-related overheads, including communication for security metadata management and secure startup costs. This paper proposes SecNPU, a CPU-decoupled and LLM-inference-optimized NPU TEE. SecNPU effectively eliminates communication overhead caused by coupled security metadata and leverages the characteristics of LLM inference to conceal security initialization latency. Experimental evaluations demonstrate that our design achieves$1.51 \times$overall secure inference speedup and$1.61 \times$secure boot performance improvement, requiring merely 1.63 % additional area and 6.6 % more power.
Xuanyao Peng, Yinghao Yang 0001, Shangjie Pan, Yujun Liang, Fengwei Zhang, Xiaowei Li 0001
ICCD2
2025 Athena: Accelerating Quantized Convolutional Neural Networks under Fully Homomorphic Encryption
abstract
Deep learning under FHE is difficult due to two aspects: (1) formidable amount of ciphertext computations like convolutions, so frequent bootstrapping is inevitable which in turn exacerbates the problem; (2) lack of the support to various non-linear functions in terms of the diversity and accuracy.Previous work primarily used the CKKS-based approach, which requires large parameters and places a heavy burden on the hardware.In this paper, we propose Athena, including a novel framework targeting quantized convolutional neural networks under FHE, and a specialized accelerator to release the maximum potential of the framework.Unlike the classic CKKS-based approach, Athena only requires much smaller parameters, i.e., 2 15 degree and approximately 5 MB ciphertext size.Athena uses a uniform representation, functional bootstrapping, to accurately support any type of activation functions, and is not limited to polynomial approximate fitted functions such as ReLU and sigmoid.We highlight the following results: (1) the accuracy varies by +0.01%/-0.24% compared with the plaintext quantized CNN;(2) the inference performance on the Athena accelerator achieves a speedup of 1.5× to 2.3×, an EDAP improvement of 3.8× to 9.9×, compared with state-of-the-art FHE accelerators.
Yinghao Yang 0001, Xicheng Xu, Liang Chang 0002, Xiaowei Li 0001
MICRO1
2025 Trident: The Acceleration Architecture for High-Performance Private Set Intersection
abstract
Private Set Intersection (PSI) is imperative in discovering the properties of the same data owned by two competitive parties, without revealing anything else of their respective data asset. Existing PSI solutions such as APSI and ORI-PSI suffer from severe communication and computation overhead due to inefficient communication and FHE polynomial evaluation, which hinders their deployment in practice. This issue is evident in both the upper-level protocol and the lower-level hardware platform. In this paper, we propose a novel software/hardware co-design acceleration architecture for PSI, termed as “Trident”, which includes two tightly coupled segments: from the protocol perspective, we investigate existing bottlenecks and propose a new PSI protocol with significantly less communication and computation under the security guarantee; besides, we re-architect the hardware platform by designing a PSI-specific accelerator, implemented with both FPGA and ASIC, targeting the key operations in the proposed protocol. We build a real-world experimental environment with two instantiated parties to verify the acceleration architecture, and highlight the following results: (1) up to 130$\boldsymbol{\times}$/145$\boldsymbol{\times}$speedup for the computation ofreceiverandsenderparties; (2) up to 37$\boldsymbol{\times}$reduction of communication overhead. (3) up to 93,651$\boldsymbol{\times}$and 74,326$\boldsymbol{\times}$higher energy efficiency over the CPU-based ORI-PSI and APSI, respectively.
Jinkai Zhang, Yinghao Yang 0001, Zhe Zhou 0003, Zhicheng Hu, Xin Zhao 0044, Liang Chang 0002, Xiaowei Li 0001
IEEE Trans. Computers2
2023 Poseidon: Practical Homomorphic Encryption Accelerator
abstract
With the development of the important solution for privacy computing, the explosion of data size and computing intensity in Fully Homomorphic Encryption (FHE) has brought enormous challenges to the hardware design. In this paper, we propose a practical FHE accelerator - "Poseidon", which focuses on improving the hardware resource and bandwidth consumption. Poseidon supports complex FHE operations like Bootstrapping, Keyswitch, Rotation and so on, under limited FPGA resources. It refines these operations by abstracting five key operators: Modular Addition (MA), Modular Multiplication (MM), Number Theoretic Transformation (NTT), Automorphsim and Shared Barret Reduction (SBT). These operators are combined and reused to implement higher-level FHE operations. To utilize the FPGA resources more efficiently and improve the parallelism, we adopt the radix-based NTT algorithm and propose HFAuto, an optimized automorphism implementation suitable for FPGA. Then, we design the hardware accelerator based on the optimized key operators and HBM to maximize computational efficiency. We evaluate Poseidon with four domain-specific FHE benchmarks on Xilinx Alveo U280 FPGA. Empirical results show that the efficient reuse of the operator cores and on-chip storage enables superior performance compared with the state-of-the-art GPU, FPGA and accelerator ASICs. We highlight the following results: (1) up to 370× speedup over CPU for the basic operations of FHE; (2) up to 1300×/52× speedup over CPU and the FPGA solution for the key operators; (3) up to 10.6×/8.7× speedup over GPU and the ASIC solution for the FHE benchmark.
Yinghao Yang 0001, Huaizhi Zhang, Shengyu Fan, Mingzhe Zhang 0005, Xiaowei Li 0001
HPCA1
2023 Poseidon-NDP: Practical Fully Homomorphic Encryption Accelerator Based on Near Data Processing Architecture
abstract
With the development of the important solution for privacy computing—fully homomorphic encryption (FHE), the explosion of data size, and computing intensity in FHE applications brings enormous challenges to the hardware design. In this article, we propose a novel co-design scheme for FHE acceleration named “Poseidon-NDP,” which focuses on improving the efficiency of the hardware resource and the bandwidth. Specifically, we investigate the special implications of the hardware imposed by the FHE applications. It empirically shows that the FHE performance is suffered from both the intractable data movement and the computation bottleneck. Besides, we also introduce the opportunity and the challenges of accelerating FHE on near data processing (NDP) architecture. Based on such analysis, we propose an optimized technique called “NTT-fusion” to simplify the FHE operator and reduce its hardware overhead. Then, we design the accelerator based on the simplified operator to achieve maximized data and computation parallelism with limited hardware resources. Additionally, we evaluate Poseidon-NDP with 4 domain-specific FHE applications on the SmartSSD, which is a practical NDP device. The empirical studies show that the efficient co-design enables Poseidon-NDP vastly superior to the state-of-the-art FHE acceleration techniques: 1) up to$217\times /84\times $speedup over CPU and high-performance GPUs for the number theoretic transform; 2) up to$3.7\times /29\times $higher-speedup/energy delay product (EDP) over the SOTA FPGA accelerator for the FHE applications; and 3) up to$4.9\times $higher-bandwidth utilization over CPU due to the NDP-based architecture.
Yinghao Yang 0001, Xiaowei Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1