Yibo Du

dblp:260/4248 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
9since 2021 · last 2026
0009-0002-5672-676XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 7 first-author · 9 since 2021Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2026 FHEx: Transforming Generic Compute Chips into Secure FHE Engines via a Hardware-software Co-designed Framework
abstract
Fully Homomorphic Encryption (FHE) is a powerful privacy-preserving technology enabling secure computation on encrypted data, but it suffers from substantial performance overheads. Running FHE efficiently typically requires developing dedicated FHE accelerators, which can be costly and inflexible. Instead of pursuing entirely new accelerators, this paper explores an alternative paradigm: augmenting generic computing devices with a modular FHE-specific hardware extension unit (HEU) to create an efficient FHE engine. To make this paradigm viable, we propose three key innovations: (1) recognizing that some FHE operators are data-intensive and involve a massive volume of ciphertexts, we design the HEU with a 3D stacked memory-based architecture to handle data-intensive operators.We also provide software-level support to facilitate deploying FHE tasks on this extension-based architecture. (2) To capitalize on the hardware parallelism, we propose an adaptive offloading algorithm that intelligently distributes FHE operators between the computing device and the HEU. (3) To optimize the data layout and minimize the inter-tile data communications in the novel 3D stack memory, we propose a dedicated ciphertext mapping mechanism. Experimental results demonstrate that our work achieves substantial acceleration in FHE tasks.
Yibo Du, Ying Wang 0001, Mengdi Wang 0004, Cangyuan Li, Hui Li 0006, Kai Zhang 0016, Yinhe Han 0001
DATE1
2026 AutoFHE: An Automatic Hardware Generation Framework for Domain-Specific FHE Accelerators
Yibo Du, Cangyuan Li, Bing Li 0017, Mengdi Wang 0004, Yinhe Han 0001
ISCA1
2026 Unlocking Pipeline Parallelism for Bootstrapping: A Pipelined Multi-Chiplet TFHE Accelerator
Yibo Du, Mengdi Wang 0004, Cangyuan Li, Yinhe Han 0001, Ying Wang 0001
ISCA1
2026 Chiplever: A Hardware-Software Co-Design Framework Toward Extension of Chiplet System for Fully Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) is a promising privacy-preserving technique that has drawn increasing attention from academia and industry. It allows computation directly on encrypted data without decryption. However, FHE incurs intensive computations. Chiplet-based designs integrate multiple processors, delivering high performance and thereby are embraced by computation-intensive FHE tasks. Despite the chiplet-based system with various processors, it is designed for unencrypted applications, falling short in handling FHE with unique ciphertext manipulations. One common approach to make it capable of FHE is developing a new FHE accelerator. However, this approach overlooks existing abundant resources already in the system and introduces a large area overhead. In this paper, we propose Chiplever, a framework that empowers a non-FHE-tailored system to efficiently support FHE tasks via a hardware extension. Chiplever aims to leverage the existing resources already in the room for FHE tasks. To achieve this, (1) Chiplever introduces a hardware extension with an FHE unit providing efficient function support for FHE operators. (2) Chiplever proposes an FHE coordinator in the extension, which enables direct ciphertext transfer between the newly introduced extension and existing chiplets, achieving efficient integration of the extension. (3) Chiplever lowers the high-level homomorphic operations to primitive operators that can be matched by existing chiplets and constructs a fine-grained computation graph. Based on this, Chiplever employs a task scheduling algorithm, which partitions the FHE task across the extension and existing chiplets to exploit the parallelism between them and reduce the ciphertext communication overheads. With these hardware and software optimizations, Chiplever achieves efficient FHE acceleration. Compared with prior FHE ASICs, Chiplever achieves 9.6× 15.9× speedup and 6.2× 67.4× throughput improvement on TFHE, while consuming only 18.8% 35.6% of the area overhead of dedicated FHE ASICs.
Yibo Du, Ying Wang 0001, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 High-Parallel In-Memory NTT Engine with Hierarchical Structure and Even-Odd Data Mapping
abstract
The Number Theoretic Transform (NTT) significantly impacts the execution time of Fully Homomorphic Encryption (FHE) in practical applications, driving research into accelerated NTT methods. Computing-in-Memory (CIM) offers a promising solution to handle NTT's memory bottlenecks, yet efficiently implementing a CIM-based NTT engine remains challenging due to unique operations and large data sizes. We propose HP-CIM, a high-parallelism digital SRAM-based CIM NTT engine designed for large-scale NTT. HP-CIM integrates MVM-based NTT with a hierarchical SRAM architecture and novel even-odd data mapping, achieving nearly 3.08× faster execution and 4.96× energy savings compared to prior CIM-based designs.
Bing Li 0017, Huaijun Liu, Yibo Du, Ying Wang 0001
ASP-DAC3
2025 RTPU: Unifying Non-Private and Private Inference with Reconfigurable Architecture
abstract
With the rise of fully homomorphic encryption-based private inference, data centers are anticipated to simultaneously handle two disparate computational demands: plaintext-based non-private inference (NPI) and ciphertext-based private inference (PI). Unfortunately, current solutions face challenges in addressing this trend. They either depend on costly, inflexible dedicated accelerators or utilize general-purpose hardware with inferior performance. This limitation underscores the urgent need for a unified architecture capable of serving both normal and privacy-sensitive users with high efficiency.However, the fundamental disparities in computation patterns and resource management between NPI and PI make their architectural fusion intricate. To bridge this gap, we explore their inherent similarities and apply fine-grained reconfiguration to maximize resource sharing. We propose RTPU, a reconfigurable multi-core architecture that can seamlessly switch between tensor-based plaintext and polynomial ring-based ciphertext computations. Building upon its reconfigurable computing fabric and parallelization mechanism, we introduce a kernel group-based scheduling strategy to optimize hardware utilization and QoS. Experimental results show that: i) The RTPU architecture achieves near-ASIC performance and beyond-ASIC flexibility with substantial silicon reuse between NPI and PI. ii) The RTPU scheduler sustains high resource utilization for multi-tenant workloads with varying privacy requirements.
Fuping Li, Ying Wang 0001, Yinghao Yang 0001, Yibo Du, Huawei Li 0001, Yinhe Han 0001, Xiaowei Li 0001
ICCAD5
2024 Chiplever: Towards Effortless Extension of Chiplet-based System for FHE
abstract
Fully Homomorphic Encryption (FHE) is one of the most promising privacy-preserving techniques that has drawn increasing attention from academia and industry due to its ideal security. Chiplet-based designs integrate multiple dies into the package delivering high performance and thereby are embraced by the resources-hungry FHE. Despite the chiplet-based system with various specialized accelerators, it falls short in supporting FHE with the novel polynomial operations. For a chiplet-based system that is not tailored for FHE, one common approach to support FHE is designing a new dedicated accelerator, However, this full design-and-build approach overlooks the existing abundant resources of accelerators in the system and incurs repeated customization and resource waste.
Yibo Du, Ying Wang 0001, Bing Li 0017, Fuping Li, Shengwen Liang, Huawei Li 0001, Xiaowei Li 0001, Yinhe Han 0001
DAC1
2024 GPACE: An Energy-Efficient PQ-Based GCN Accelerator with Redundancy Reduction
abstract
Graph convolutional network (GCN) has been proven powerful in various tasks for it combines both neural networks and graph processing operators. However, this characteristic makes GCN exhibit hybrid execution patterns, which is unfavorable for CPUs and GPUs. Therefore, designing specialized GCN accelerators is becoming a prevalent paradigm. Unfortunately, as graph scale continues to grow, existing GCN accelerators suffer from significant bandwidth consumption and memory footprint as they neglect the inherent semantic redundancy of vertex features. Although applying Product Quantization to GCN is a promising solution to reduce the sizeable graph data via distilling semantic redundancy, it introduces novel operations with unique patterns that existing GCN accelerators cannot support. In this paper, we propose GPACE, an energy-efficient GCN accelerator that can fully harness the potential of PQ to reduce bandwidth consumption and data movement. GPACE is designed with a lookup-efficient architecture and well-optimized dataflow to support the unique data access and computation pattern of PQ-GCN. In addition to leveraging PQ to distill semantic redundancy, we exploit the operation redundancy and propose a redundancy-aware architecture to detect and reduce types of redundant operations to achieve higher energy efficiency. Evaluations show GPACE achieves high speedup and energy saving compared with CPU, GPU, and specialized GCN accelerators.
Yibo Du, Shengwen Liang, Ying Wang 0001, Huawei Li 0001, Xiaowei Li 0001, Yinhe Han 0001
DATE1
2023 PANG: A Pattern-Aware GCN Accelerator for Universal Graphs
abstract
Graph convolutional neural network (GCN) extends deep learning to process graph data and demonstrates superior performance. However, due to the irregularity, graphs show inconsistent patterns across different regions, which leads to distinctions in data reusability and edge processing activity, and consequently poses impacts on hardware efficiency and resource utility. Prior accelerators seldom explore the distinct patterns across graph regions and adopt a fixed strategy for the whole graph without consideration for region-specific characteristics. In this paper, we identify the inconsistent patterns of graphs and characterize the distinctions between graph regions. Then, we propose an adaptive dataflow to adapt the region-specific patterns. Third, we implement PANG, a pattern-aware accelerator that can dynamically adjust the dataflow to exploit the reusability and alleviate the frequent destination switching. Evaluated on real-world datasets, PANG achieves significant performance improvement.
Yibo Du, Ying Wang 0001, Shengwen Liang, Huawei Li 0001, Xiaowei Li 0001, Yinhe Han 0001
ICCD1