Fahong Zhang 0004

dblp:389/7237 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0008-8613-5323ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 57% Memory systems · 36% Integrated circuit design · 6%
Network and information security
2 papers
Cryptographic primitives and cryptanalysis · 63% Privacy and data protection · 18% Cryptographic protocols and secure computation · 18%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
cryptographic accelerator
1.622025
Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator · ACM Trans. Archit. Code Optim. 2025
A High-Throughput Private Inference Engine Based on 3D Stacked Memory · DAC 2024
Cryptographic primitives and cryptanalysis › homomorphic encryption
fully homomorphic encryption
0.912025
Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator · ACM Trans. Archit. Code Optim. 2025
Cryptographic primitives and cryptanalysis › homomorphic encryption › fully homomorphic encryption › TFHE
programmable bootstrapping
0.912025
Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator · ACM Trans. Archit. Code Optim. 2025
Cryptographic primitives and cryptanalysis › homomorphic encryption › fully homomorphic encryption
TFHE
0.912025
Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator · ACM Trans. Archit. Code Optim. 2025
Privacy and data protection
privacy-preserving computation
0.812024
A High-Throughput Private Inference Engine Based on 3D Stacked Memory · DAC 2024
Cryptographic protocols and secure computation
secure inference
0.812024
A High-Throughput Private Inference Engine Based on 3D Stacked Memory · DAC 2024
Memory systems
3d-stacked memory
0.812024
A High-Throughput Private Inference Engine Based on 3D Stacked Memory · DAC 2024
Memory systems › DRAM › DRAM architecture
embedded DRAM
0.812024
A High-Throughput Private Inference Engine Based on 3D Stacked Memory · DAC 2024
Hardware accelerators and domain-specific architectures › cryptographic accelerator
fully homomorphic encryption accelerator
0.812024
A High-Throughput Private Inference Engine Based on 3D Stacked Memory · DAC 2024
Integrated circuit design
ASIC design
0.312025
Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator · ACM Trans. Archit. Code Optim. 2025

Methods — techniques the papers use, named apart from their topics

special-prime processing element · 1.7hybrid dataflow · 1.7software-hardware co-design · 1.5polynomial decomposition · 1.53d hybrid bonding · 1.5
YearPublicationVenuePosition
2025 Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator
abstract
Fully homomorphic encryption over torus (TFHE) enables the execution of arbitrary functions on encrypted data through programmable bootstrapping (PBS). However, performing all operations on ciphertext during PBS results in high computational and memory requirements, limiting the deployment of PBS in real-world scenarios. Previous TFHE accelerator designs have attempted to improve performance by employing specific dataflow and functional units, but these techniques may require large off-chip bandwidth or on-chip storage when scaling up computation capacity. Additionally, the design of specialized functional units may limit the utilization of computation units when facing dynamic secure parameter settings. To address these challenges and further improve PBS throughput in TFHE, we propose Matrix , an ASIC-based architecture that balances off-chip bandwidth and on-chip storage according to the execution flow of PBS. In Matrix , we utilize a unified special-prime-based processing element (PE) that achieves high utilization with minimal resource overhead. Furthermore, we propose a hybrid PBS dataflow that can efficiently reduce computation complexity and memory requirements. Compared to state-of-the-art TFHE accelerators, Matrix achieves 1.43 × -5.66 × throughput improvement for PBS. For ZAMA Deep-NN benchmark, we achieve 525.60× and 68.06× speedup compared to CPU and GPU, respectively. 1
Ling Liang 0003, Fahong Zhang 0004, Zhirui Li, Xin Fan 0009, Dimin Niu, Meng Li 0004, Zhiyong Li 0016, Zongwei Wang 0001, Hongzhong Zheng, Yimao Cai, Yuan Xie 0001
ACM Trans. Archit. Code Optim.3
2024 A High-Throughput Private Inference Engine Based on 3D Stacked Memory
abstract
Fully Homomorphic Encryption (FHE) enables unlimited computation depth, allowing privacy-enhanced neural network inference tasks directly on the ciphertext. However, existing FHE architectures suffer from the memory access bottleneck. This work proposes a High-throughput FHE engine for private inference (PI) based on 3D stacked memory (H3). H3 adopts the software-hardware co-design that dynamically adjusts the polynomial decomposition during the PI process to minimize the computation and storage overhead at a fine granularity. With 3D hybrid bonding, H3 integrates a logic die with a multi-layer embedded DRAM, routing data efficiently to the processing unit array through an efficient broadcast mechanism. H3 consumes 192mm2 when implemented using a 28nm logic process. It achieves 1.36 million LeNet-5 or 920 ResNet-20 PI per minute, surpassing existing 7nm accelerators by 52%. This demonstrates that 3D memory is a promising technology to promote the performance of FHE.
Ling Liang 0003, Zhirui Li, Fahong Zhang 0004, Yanheng Lu
DAC5