Haoqi He

dblp:368/7607 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0009-8270-4442ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HiQ-Lip: A Hierarchical Quantum-Classical Method for Global Lipschitz Constant Estimation of ReLU Networks
abstract
Estimating the global Lipschitz constant of neural networks is crucial for understanding and improving their robustness and generalization capabilities. However, precise calculations are NP-hard, and current semidefinite programming (SDP) methods face challenges such as high memory usage and slow processing speeds. In this paper, we propose HiQ-Lip, a hybrid quantum-classical hierarchical method that leverages Coherent Ising Machines (CIMs) to estimate the global Lipschitz constant. We tackle the estimation by converting it into a Quadratic Unconstrained Binary Optimization (QUBO) problem and implement a multilevel graph coarsening and refinement strategy to adapt to the constraints of contemporary quantum hardware. Our experimental evaluations on fully connected neural networks demonstrate that HiQ-Lip not only provides estimates comparable to state-of-the-art methods but also significantly accelerates the computation process. In specific tests involving two-layer neural networks with 256 hidden neurons, HiQ-Lip doubles the solving speed and offers more accurate upper bounds than the existing best method, LiPopt. These findings highlight the promising utility of small-scale quantum devices in advancing the estimation of neural network robustness.
Haoqi He, Wenzhi Xu, Ruoying Liu, Xiaokai Lin, Kai Wen
AAAI1
2026 Maverick: Rethinking TFHE Bootstrapping on GPUs via Algorithm-Hardware Co-Design
abstract
Fully homomorphic encryption (FHE) enables arbitrary computation over encrypted data (ciphertext) without compromising confidentiality. Within this family, TFHE features versatile bootstrapping mechanisms that is attractive for security-critical applications. However, its prohibitive computational cost severely limits practical deployment. While hardware acceleration is promising, mere compute scaling fails to overcome the inherent barriers. In particular, the combination of limited algorithmic parallelism and inadequate understanding of hardware behaviors prevents full exploitation of the available performance headroom.
Haoqi He, Lutan Zhao, Qingyun Niu, Dan Meng 0002, Rui Hou 0001
ASPLOS (2)2
2026 HERDS: Multi-key Fully Homomorphic Encryption with Sublinear Bootstrapping
Binwu Xiang, Seonhong Min, Intak Hwang, Haoqi He, Yuanju Wei, Kang Yang 0002, Jiang Zhang 0001, Yi Deng 0002, Yu Yu 0001
EUROCRYPT (4)5
2026 Peregrine: Accelerating TFHE Bootstrapping on GPUs via Multi-Level External Product Co-Design
abstract
Fully Homomorphic Encryption (FHE) is a ground-breaking cryptographic technology that enables computation directly on encrypted data, but its practical adoption continues to be hindered by high computational costs. GPUs have emerged as an increasingly attractive acceleration platform, offering massive parallelism and architectural flexibility to accommodate rapidly evolving FHE algorithms. Despite notable advances, most efforts remain confined to isolated, single-level optimizations, limiting their ability to push performance boundaries. In this paper, we present Peregrine, an efficient GPU-based TFHE acceleration built upon multi-level co-design of external product (EP) operations across parallelism, implementation, and scheduling. First, we propose a synchronization-free key unrolling technique that restructures the execution pipeline via operator decoupling, thereby unlocking greater EP-level parallelism. Second, we consolidate fragmented operators into a matrix-centric execution pattern, yielding a high-arithmetic-intensity kernel that substantially enhances external product efficiency. Third, we propose a hierarchical tiling strategy that reformats ring ciphertexts into the Module structure and schedules polynomiallevel tiles for fine-grained GPU mapping of EP operations. Experimental results show that Peregrine outperforms up to$176.9 \times$and$2.2 \times$over state-of-the-art CPU and GPU baselines, respectively, demonstrating its strong applicability to security-critical workloads.
Haoqi He, Lutan Zhao, Dan Meng 0002, Rui Hou 0001
HPCA1
2025 ShiftPIR: An Efficient PIR System with Gravity Shifting from Client to Server
abstract
We present ShiftPIR, a single-server Private Information Retrieval (PIR) protocol that gravity shifts both computation and communication overhead from the client to the server, thereby significantly improving overall efficiency. This shift is driven by the growing asymmetry between resource-constrained clients and compute-intensive servers, where server-side tasks can be effectively parallelized and scaled. To achieve this, ShiftPIR introduces a novel request generation method in which the client transmits only a compact plaintext offset derived from pre-uploaded seed ciphertexts. The server then reconstructs the full query ciphertexts using homomorphic rotations, eliminating the need for costly ciphertext generation and transmission on the client side. We further design a highly parallelizable query expansion mechanism that removes data dependencies between ciphertext rotations, enabling efficient GPU-based execution. Our experiments demonstrate that ShiftPIR reduces client-side latency to microseconds while maintaining a communication cost within 4X of the non-private baseline—far outperforming prior protocols with 104 -105 × overhead. Compared to the state-of-the-art protocol YPIR, ShiftPIR achieves up to 26X lower end-to-end latency.
Lutan Zhao, Haoqi He, Wenzhe Lv, Dan Meng 0002, Rui Hou 0001
CCS5
2025 Q-Detection: A Quantum-Classical Hybrid Poisoning Attack Detection Method
abstract
Data poisoning attacks pose significant threats to machine learning models by introducing malicious data into the training process, thereby degrading model performance or manipulating predictions. Detecting and sifting out poisoned data is an important method to prevent data poisoning attacks. Limited by classical computation frameworks, upcoming larger-scale and more complex datasets may pose difficulties for detection. We introduce the unique speedup of quantum computing for the first time in the task of detecting data poisoning. We present Q-Detection, a quantum-classical hybrid defense method for detecting poisoning attacks. Q-Detection also introduces the Quantum Weight-Assigning Network, which is optimized using quantum computing devices. Experimental results using multiple quantum simulation libraries show that Q-Detection effectively defends against label manipulation and backdoor attacks. The metrics demonstrate that Q-Detection consistently outperforms the baseline methods and is comparable to the state-of-the-art. Theoretical analysis shows that Q-Detection is expected to achieve more than a 20% speedup using quantum computing power.
Haoqi He, Xiaokai Lin, Jiancai Chen
IJCAI1
2025 Chameleon: An Efficient FHE Scheme Switching Acceleration on GPUs
abstract
Fully homomorphic encryption (FHE) enables direct computation on encrypted data, making it a crucial technology for privacy protection. However, FHE suffers from significant performance bottlenecks. In this context, GPU acceleration offers a promising solution to bridge the performance gap. Existing efforts primarily focus on single-class FHE schemes, which fail to meet the diverse requirements of data types and functions, prompting the development of hybrid multi-class FHE schemes. However, studies have yet to thoroughly investigate specific GPU optimizations for hybrid FHE schemes. In this paper, we present an efficient GPU-based FHE scheme switching acceleration named Chameleon. First, we propose a scalable NTT acceleration design that adapts to larger CKKS polynomials and smaller TFHE polynomials. Specifically, Chameleon tackles synchronization issues by fusing stages to reduce synchronization, employing polynomial coefficient shuffling to minimize synchronization scale, and utilizing an SM-aware combination strategy to identify the optimal switching point. Second, Chameleon is the first to comprehensively analyze and optimize critical switching operations. It introduces CMux-level parallelization to accelerate LUT evaluation and a homomorphic rotation-free matrixvector multiplication to improve repacking efficiency. Finally, Chameleon outperforms the state-of-the-art GPU implementations by 1.23× in CKKS HMUL and 1.15× in bootstrapping. It also achieves up to 4.87× and 1.51× speedups for TFHE bootstrapping compared to CPU and GPU versions, respectively, and delivers a 67.3× average speedup for scheme switching over CPU-based implementation.
Haoqi He, Lutan Zhao, Peinan Li, Zhihao Li 0001, Dan Meng 0002, Rui Hou 0001
IEEE Trans. Parallel Distributed Syst.2
2024 Quantum Annealing and GNN for Solving TSP with QUBO
Haoqi He
AAIM (2)1
2024 Quantum Lung Segmentation: QCU-Net Applied to Chest X-Ray Images
Haoqi He, Mingkai Huang
AAIM (2)1