Qingyun Niu

dblp:41/8340 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0005-9801-2417ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 70% GPUs and heterogeneous computing · 23% Processor architecture and microarchitecture · 7%
Network and information security
2 papers
Cryptographic primitives and cryptanalysis · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic primitives and cryptanalysis › homomorphic encryption
fully homomorphic encryption
2.022026
UniFHE: Faster Accelerator for FHE with Diverse Algebraic Structure and Balanced Memory System · HPCA 2026
Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-Design · ASPLOS (2) 2026
Hardware accelerators and domain-specific architectures
cryptographic accelerator
2.022026
UniFHE: Faster Accelerator for FHE with Diverse Algebraic Structure and Balanced Memory System · HPCA 2026
Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-Design · ASPLOS (2) 2026
Cryptographic primitives and cryptanalysis › homomorphic encryption › fully homomorphic encryption
TFHE bootstrapping
1.012026
Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-Design · ASPLOS (2) 2026
Hardware accelerators and domain-specific architectures › cryptographic accelerator
fully homomorphic encryption accelerator
1.012026
UniFHE: Faster Accelerator for FHE with Diverse Algebraic Structure and Balanced Memory System · HPCA 2026
GPUs and heterogeneous computing
GPU computing
1.012026
Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-Design · ASPLOS (2) 2026
Cryptographic primitives and cryptanalysis
homomorphic encryption
0.312026
Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-Design · ASPLOS (2) 2026
Processor architecture and microarchitecture › arithmetic unit
arithmetic unit design
0.312026
UniFHE: Faster Accelerator for FHE with Diverse Algebraic Structure and Balanced Memory System · HPCA 2026

Methods — techniques the papers use, named apart from their topics

on-chip plaintext encoding · 2.0multi-pipeline architecture · 2.0general arithmetic unit · 2.0algorithm-hardware co-design · 2.0
YearPublicationVenuePosition
2026 Sub-Millisecond Gate Bootstrapping
abstract
Gate bootstrapping is a core primitive that enables arbitrary circuit evaluation in fully homomorphic encryption (FHE), where blind rotation remains the dominant performance bottleneck. In this work, we present a sub-millisecond NTRU-based gate bootstrapping scheme that achieves state-of-the-art performance through coordinated algorithmic, software, and hardware-level optimizations.
Chunling Chen, Zhihao Li 0001, Qingyun Niu, Xianhui Lu, Ruida Wang, Lutan Zhao, Rui Hou 0001
AsiaCCS3
2026 Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-Design
abstract
Fully homomorphic encryption (FHE) enables arbitrary computation over encrypted data (ciphertext) without compromising confidentiality. Within this family, TFHE features versatile bootstrapping mechanisms that is attractive for security-critical applications. However, its prohibitive computational cost severely limits practical deployment. While hardware acceleration is promising, mere compute scaling fails to overcome the inherent barriers. In particular, the combination of limited algorithmic parallelism and inadequate understanding of hardware behaviors prevents full exploitation of the available performance headroom.
Haoqi He, Lutan Zhao, Qingyun Niu, Dan Meng 0002, Rui Hou 0001
ASPLOS (2)4
2026 Thunder: Efficient Multi-node FHE Acceleration Framework via In-Transit Computation
Lutan Zhao, Qingyun Niu, Yinhang Zheng, Zhengbang Yang, Boyan Zhao, Rui Hou 0001
Euro-Par (1)3
2026 UniFHE: Faster Accelerator for FHE with Diverse Algebraic Structure and Balanced Memory System
abstract
Fully homomorphic encryption (FHE) enables computations on encrypted data. Existing FHE schemes are primarily categorized into RLWE-based word-wise schemes and LWEbased bit-wise schemes. Efficient combination of different FHE schemes adapted to real-world applications has emerged as a research focus. This paper proposes UniFHE, the first FHE accelerator that supports diverse algebraic structures using general arithmetic units to achieve higher performance. UniFHE is compatible with both RLWE-based and LWE-based FHE schemes without modifications to their original algorithmic designs. To support both finite ring and complex field operations, UniFHE introduces a general arithmetic unit and further constructs core computation structures. To balance on-chip memory demands across different schemes, UniFHE adopts a multi-pipeline architecture for LWE-based schemes. The core functional units for RLWE-based schemes are spliced based on the LWE-based pipelines. Furthermore, an on-chip plaintext encoding mechanism significantly reduces off-chip memory bandwidth demands. Experimental results show that, beyond superior area and energy efficiency, UniFHE delivers up to$13.6 \times$higher performance compared to scheme-specific accelerator combinations. Moreover, in hybrid schemes, UniFHE achieves a$3.2 \times$speedup over the state-of-the-art unified FHE accelerator Trinity.
Qingyun Niu, Lutan Zhao, Dan Meng 0002, Rui Hou 0001
HPCA1
2026 A survey of optimization techniques for bootstrapping algorithms in FHE
abstract
Abstract Fully Homomorphic Encryption (FHE) enables arbitrary computation on encrypted data without decryption, making it a cornerstone of privacy-preserving outsourcing, such as cloud computing. However, homomorphic operations cause ciphertext noise to grow until decryption fails. The efficient solution is bootstrapping, which refreshes the noise in FHE ciphertexts to sustain arbitrary deep homomorphic evaluation. But in practice, bootstrapping consumes over 50% of total execution time, posing a serious obstacle to FHE adoption. This paper presents a systematic survey of FHE bootstrapping algorithms and their optimizations. We organize existing works into three main paradigms: word-wise bootstrapping for BGV, BFV, and CKKS schemes; bit-wise bootstrapping for FHEW and TFHE schemes; and hybrid bootstrapping, which leverages both word-wise schemes and bit-wise schemes. We analyze the evolution of crucial techniques, highlight latest advances in reducing latency, enhancing parallelism, and controlling noise growth, and compare the advantages and limitations of different schemes. Finally, we discuss emerging research trends.
Lutan Zhao, Ruida Wang, Qingyun Niu, Xianhui Lu, Dan Meng 0002, Rui Hou 0001
Cybersecur.5