EDBT 2026 Demo / reviewers in the wild / expert
Qian Lou
dblp:207/3962
· DBLP profile ↗
50ranked-venue papers
13as first author
39since 2021 · last 2026
0000-0001-5462-2567ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 10 first-author · 22 since 2021Systems, architecture and hardware · 16 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Security and privacy · 5 · 5 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conjunctive Prompt Attacks in Multi-Agent LLM SystemsabstractMost LLM safety work studies single-agent models, but many real applications rely on multiple interacting agents.In these systems, prompt segmentation and inter-agent routing create attack surfaces that single-agent evaluations miss.We study conjunctive prompt attacks, where a trigger key in the user query and a hidden adversarial template in one compromised remote agent each appear benign alone but activate harmful behavior when routing brings them together.We consider an attacker who changes neither model weights nor the client agent and instead controls only trigger placement and template insertion.Across star, chain, and DAG topologies, routing-aware optimization substantially increases attack success over non-optimized baselines while keeping false activations low.Existing defenses, including PromptGuard, Llama-Guard variants, and system-level controls such as tool restrictions, do not reliably stop the attack because no single component appears malicious in isolation.These results expose a structural vulnerability in agentic LLM pipelines and motivate defenses that reason over routing and cross-agent composition. Nokimul Hasan Arif, Qian Lou, Mengxin Zheng |
ACL (1) | 2 |
| 2026 | ReliaFHE: Resilient Design for Fully Homomorphic Encryption AcceleratorsabstractThe significant computational complexity of Fully Homomorphic Encryption (FHE) has prompted numerous accelerator designs. However, existing FHE accelerators often implicitly assume that all computations are executed reliably, overlooking the fact that modern FHE schemes can be highly vulnerable to hardware faults: even a single-bit error in a ciphertext can cascade into widespread plaintext corruption. Ruizhi Zhu, Mengxin Zheng, Qian Lou, Xin Xin 0008 |
ASPLOS (2) | 5 |
| 2026 | HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECC
Ruizhi Zhu, Yanan Guo 0002, Huize Li, Weidong Cao 0001, Qian Lou, Xin Xin 0008 |
ISCA | 5 |
| 2026 | Heterogeneous Multi-Agent Reinforcement Learning with Attention for Cooperative and Scalable Feature TransformationabstractFeature transformation enhances downstream task performance by generating informative features through mathematical feature crossing. Despite the advancements in deep learning, feature transformation remains essential, particularly for structured data, where deep models often struggle to capture complex feature interactions effectively. Prior literature on automated feature transformation has achieved notable success but often relies on heuristics or exhaustive searches, leading to inefficient and time-consuming processes. Recent works employ reinforcement learning (RL) to enhance traditional approaches through a more effective trial-and-error way. However, two key limitations remain: 1) Dynamic feature expansion during the transformation process, which introduces instability and increases the time complexity of the learning procedure for RL agents; 2) Insufficient cooperation and communication between agents, which results in suboptimal feature crossing operations and degraded model performance. To address them, we propose a novel heterogeneous multi-agent RL framework to enable cooperative and scalable feature transformation. The framework comprises three heterogeneous agents, grouped into two types, each designed to select essential features and operations for feature crossing. To enhance communication among these agents, we implement a shared critic mechanism that facilitates information exchange during the feature transformation process. This collaboration enables the agents to learn more intelligent and effective transformation policies. To handle the dynamically expanding feature space, we tailor multi-head attention-based feature agents to select suitable features for feature crossing. This design facilitates scalable decision-making and effective candidate selection based on comprehensive global feature space information. Additionally, we introduce a state encoding technique during the optimization process to stabilize and enhance the learning dynamics of the RL agents, resulting in more robust and reliable transformation policies. Finally, we conduct extensive experiments to validate the effectiveness, efficiency, robustness, and interpretability of our model. Our code and dataset are publicly available on GitHub. Tao Zhe, Huazhen Fang, Kunpeng Liu 0001, Qian Lou, Tamzidul Hoque, Dongjie Wang 0001 |
KDD (1) | 4 |
| 2026 | QNBAD: Quantum Noise-induced Backdoor Attacks against Zero Noise Extrapolation
Cheng Chu, Qian Lou, Fan Chen 0001, Lei Jiang 0001 |
NDSS | 2 |
| 2026 | Efficient Arithmetic-and-Comparison Homomorphic Encryption with Space Switching
Erwin Eko Wahyudi, Yan Solihin, Qian Lou |
SP | 3 |
| 2026 | Efficient privacy-preserving sparse matrix-vector multiplication using homomorphic encryption
Yang Gao 0001, Gang Quan, Wujie Wen, Scott Piersall, Qian Lou, Liqiang Wang 0001 |
Inf. Sci. | 5 |
| 2026 | SoK: Can Fully Homomorphic Encryption Support General AI Computation? A Functional and Cost AnalysisabstractArtificial intelligence (AI) increasingly powers sensitive applications in domains such as healthcare and finance, relying on both extit{linear operations} (e.g., matrix multiplications in large language models) and extit{non-linear operations} (e.g., sorting in retrieval-augmented generation). Fully homomorphic encryption (FHE) has emerged as a promising tool for privacy-preserving computation, but it remains unclear whether existing methods can support the full spectrum of AI workloads that combine these operations. In this SoK, we ask: extit{Can FHE support general AI computation?} We provide both a functional analysis and a cost analysis. First, we categorize ten distinct FHE approaches and evaluate their ability to support general computation. We then identify three promising candidates and benchmark workloads that mix linear and non-linear operations across different bit lengths and SIMD parallelization settings. Finally, we evaluate five real-world, privacy-sensitive AI applications that instantiate these workloads. Our results quantify the costs of achieving general computation in FHE and offer practical guidance on selecting FHE methods that best fit specific AI application requirements. Our codes are available at https://github.com/UCF-ML-Research/FHE-AI-Generality. Wei Zhang 0076, Mengxin Zheng, Minxuan Zhou, Yushun Dong, Dongjie Wang 0001, Jiafeng Xie, David Mohaisen, Hongyi Wu, Qian Lou |
Proc. Priv. Enhancing Technol. | 14 |
| 2025 | Corrosion Hammer: A Self-Activated Bit-Flip Attack to the Processing-In-Memory AcceleratorabstractIn this paper, taking ReRAM-based PIM accelerators as an example, we present a novel attack framework called Corrosion Hammer, which builds based on the Bit Flip Attack (BFA).Unlike previous BFA methods that require explicit memory fault injection techniques, such as Row Hammer, to modify sensitive bits in the victim Neural Network model, Corrosion Hammer implants the trojan during the hardware-software co-design phase and flips sensitive bits using read disturbance, which is a common noise in ReRAM caused by normal read operations.Furthermore, we explore the impact of inputs on the activation time consumption of the trojan and propose a method to expedite activation using normal input.Our experimental results demonstrate that Corrosion Hammer achieves an extremely covert trojan implantation and activation method, with an adversarial attack success rate of 92.46%.Additionally, using a specially designed method, Trojan activation is 61.98× faster compared to activation in an undisturbed normal operation state.It provides a way to significantly speed up the Trojan activation. Mengxin Zheng, Shengyu Fan, Qian Lou, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005 |
CF | 4 |
| 2025 | zkVC: Fast Zero-Knowledge Proof for Private and Verifiable ComputingabstractIn the context of cloud computing, services are held on cloud servers, where the clients send their data to the server and obtain the results returned by server. However, the computation, data and results are prone to tampering due to the vulnerabilities on the server side. Thus, verifying the integrity of computation is important in the client-server setting. The cryptographic method known as Zero-Knowledge Proof (ZKP) is renowned for facilitating private and verifiable computing. ZKP allows the client to validate that the results from the server are computed correctly without violating the privacy of the server’s intellectual property. Zero-Knowledge Succinct NonInteractive Argument of Knowledge (zkSNARKs), in particular, has been widely applied in various applications like blockchain and verifiable machine learning. Despite their popularity, existing zkSNARKs approaches remain highly computationally intensive. For instance, even basic operations like matrix multiplication require an extensive number of constraints, resulting in significant overhead. In addressing this challenge, we introduce $z k V C$, which optimizes the ZKP computation for matrix multiplication, enabling rapid proof generation on the server side and efficient verification on the client side. zkVC integrates optimized ZKP modules, such as Constraint-reduced Polynomial Circuit (CRPC) and Prefix-Sum Query (PSQ), collectively yielding a more than $\mathbf{1 2}$-fold increase in proof speed over prior methods. The code is available at https://github.com/UCF-Lou-Lab-PET/zkformer. Yancheng Zhang, Mengxin Zheng, Jingtong Hu, Lei Ju 0001, Yan Solihin, Qian Lou |
DAC | 8 |
| 2025 | CipherPrune: Efficient and Scalable Private Transformer InferenceabstractPrivate Transformer inference using cryptographic protocols offers promising solutions for privacy-preserving machine learning; however, it still faces significant runtime overhead (efficiency issues) and challenges in handling long-token inputs (scalability issues). We observe that the Transformer's operational complexity scales quadratically with the number of input tokens, making it essential to reduce the input token length. Notably, each token varies in importance, and many inputs contain redundant tokens. Additionally, prior private inference methods that rely on high-degree polynomial approximations for non-linear activations are computationally expensive. Therefore, reducing the polynomial degree for less important tokens can significantly accelerate private inference. Building on these observations, we propose \textit{CipherPrune}, an efficient and scalable private inference framework that includes a secure encrypted token pruning protocol, a polynomial reduction protocol, and corresponding Transformer network optimizations. At the protocol level, encrypted token pruning adaptively removes unimportant tokens from encrypted inputs in a progressive, layer-wise manner. Additionally, encrypted polynomial reduction assigns lower-degree polynomials to less important tokens after pruning, enhancing efficiency without decryption. At the network level, we introduce protocol-aware network optimization via a gradient-based search to maximize pruning thresholds and polynomial reduction conditions while maintaining the desired accuracy. Our experiments demonstrate that CipherPrune reduces the execution overhead of private Transformer inference by approximately $6.1\times$ for 128-token inputs and $10.6\times$ for 512-token inputs, compared to previous methods, with only a marginal drop in accuracy. The code is publicly available at https://github.com/UCF-Lou-Lab-PET/cipher-prune-inference. Yancheng Zhang, Mengxin Zheng, Mimi Xie, Mingzhe Zhang 0005, Lei Jiang 0001, Qian Lou |
ICLR | 7 |
| 2025 | DictPFL: Efficient and Private Federated Learning on Encrypted GradientsabstractFederated Learning (FL) enables collaborative model training across institutions without sharing raw data. However, gradient sharing still risks privacy leakage, such as gradient inversion attacks. Homomorphic Encryption (HE) can secure aggregation but often incurs prohibitive computational and communication overhead. Existing HE-based FL methods sit at two extremes: encrypting all gradients for full privacy at high cost, or partially encrypting gradients to save resources while exposing vulnerabilities. We present **DictPFL**, a practical framework that achieves full gradient protection with minimal overhead. DictPFL encrypts every transmitted gradient while keeping non-transmitted parameters local, preserving privacy without heavy computation. It introduces two key modules: **Decompose-for-Partial-Encrypt (DePE)**, which decomposes model weights into a static dictionary and an updatable lookup table—only the latter is encrypted and aggregated, while the static dictionary remains local and requires neither sharing nor encryption; and **Prune-for-Minimum-Encrypt (PrME)**, which applies encryption-aware pruning to minimize encrypted parameters via consistent, history-guided masks. Experiments show that DictPFL reduces communication cost by 402-748$\times$ and accelerates training by 28-65$\times$ compared to fully encrypted FL, while outperforming state-of-the-art selective encryption methods by 51-155$\times$ in overhead and 4-19$\times$ in speed. Remarkably, DictPFL’s runtime is within 2$\times$ of plaintext FL, demonstrating, for the first time, that HE-based private federated learning is practical for real-world deployment. The code is publicly available at https://github.com/UCF-ML-Research/DictPFL. Yuzhang Shang, Shangqian Gao, Rui Ning, Mengxin Zheng, Xiaoqian Jiang, Qian Lou |
NeurIPS | 8 |
| 2025 | DataSeal: Ensuring the Verifiability of Private Computation on Encrypted DataabstractFully Homomorphic Encryption (FHE) allows computations to be performed directly on encrypted data without needing to decrypt it first. This “encryption-in-use” feature is crucial for securely outsourcing computations in privacy-sensitive areas such as healthcare and finance. Nevertheless, in the context of FHE-based cloud computing, clients often worry about the integrity and accuracy of the outcomes. This concern arises from the potential for a malicious server or server-side vulnerabilities that could result in tampering with the data, computations, and results. Ensuring integrity and verifiability with low overhead remains an open problem, as prior attempts have not yet achieved this goal. To tackle this challenge and ensure the verification of FHE's private computations on encrypted data, we introduce DataSeal, which combines the low overhead of the algorithm-based fault tolerance (ABFT) technique with the confidentiality of FHE, offering high efficiency and verification capability. Through thorough testing in diverse contexts, we demonstrate that DataSeal achieves much lower overheads for providing computation verifiability for FHE than other techniques that include MAC, ZKP, and TEE. DataSeal's space and computation overheads decrease to nearly negligible as the problem size increases. Muhammad Husni Santriaji, Yancheng Zhang, Qian Lou, Yan Solihin |
SP | 4 |
| 2025 | Corrosion Hammer: a self-activated bit-flip attack to the processing-in-memory acceleratorabstractAbstract The Resistive Random-Access-Memory (ReRAM) crossbar-based Processing-In-Memory (PIM) accelerator shows great promise in accelerating neural networks (NNs). This technique boasts low energy consumption and exceptional performance in multiplication and accumulations (MAC) operations, making ReRAM-based PIM accelerators an ideal solution for intelligent computing in wearable and low-power mobile devices. However, security concerns related to PIM have not been adequately addressed. In this paper, we present a new attack framework called SolutionName for ReRAM-based PIM accelerators. SolutionName builds upon the Bit Flip Attack (BFA), a weight modification attack that manipulates the NN function by flipping specific bits in the deployed quantized NN model. Unlike previous BFA methods that require explicit memory fault injection techniques, such as Row Hammer, to modify sensitive bits in the victim NN, SolutionName implants the trojan during the hardware-software co-design phase and flips sensitive bits using read disturbance. Read disturbance is a common noise in ReRAM caused by normal read operations. This approach enables the trojan to be activated quietly during normal use, eliminating the need for explicit attacks. Furthermore, we explore the impact of inputs on the activation time of the trojan and propose a method to expedite activation using normal input. Our experimental results demonstrate that SolutionName achieves an extremely covert trojan implantation and activation method, with an adversarial attack success rate of 94.38%. Additionally, with a specially designed method, the trojan activation can be accelerated on average by 61.98 $$\times $$ × , providing controllable activation. Mengxin Zheng, Shengyu Fan, Qian Lou, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005 |
Cybersecur. | 4 |
| 2025 | DAHE: Parameter-Adaptive and Memory Efficient FPGA Acceleration of Homomorphic EncryptionabstractWhile homomorphic encryption (HE) has been well-recognized as a promising data privacy protection technique, there are many challenges to the real-world deployment of HE applications. In this work, we propose a design flow for parameter-adaptive and memory-efficient FPGA acceleration of homomorphic encryption. In the framework, we explore the correlations between HE parameter selection to meet various design objectives and the huge design space due to underlying FPGA hardware resource allocation. Particularly, we demonstrate that adaptive management of the FPGA memory hierarchy is crucial to supporting diverse cryptosystem parameter selection for application-level security, accuracy, and performance requirements. We propose a resource-efficient and flexible micro-architectural design for HE operations, where data access patterns in various pipeline execution stages are optimized for high memory bandwidth utilization. Furthermore, a memory-aware performance model is built for automatic design space exploration for cryptosystem parameter selection and hardware resource provisioning. Experimental results show 1.50X and 1.16X speedup for the NTT and Rotation operations w.r.t. the state-of-the-art FPGA implementation. Meanwhile, the proposed framework generates flexible and high-performance accelerator code for real HE application kernels with different cryptosystem parameters on a wide range of FPGA devices. Yilan Zhu, Honghui You, Wei Zhang 0173, Jiming Xu, Qian Lou, Shoumeng Yan, Lei Ju 0001 |
IEEE Trans. Computers | 5 |
| 2024 | BoostCom: Towards Efficient Universal Fully Homomorphic Encryption by Boosting the Word-wise ComparisonsabstractFully Homomorphic Encryption (FHE) allows for the execution of computations on encrypted data without the need to decrypt it first, offering significant potential for privacy-preserving computational operations. Emerging arithmetic-based FHE schemes (ar-FHE), like BGV, demonstrate even better performance in word-wise comparison operations over non-arithmetic FHE (na-FHE) schemes, such as TFHE, especially for basic tasks like comparing values, finding maximums, and minimums. This shows the universality of ar-FHE in effectively handling both arithmetic and non-arithmetic operations without the expensive conversion between arithmetic and non-arithmetic FHEs. We refer to universal arithmetic Fully Homomorphic Encryption as uFHE. The arithmetic operations in uFHE remain consistent with those in the original arithmetic FHE, which have seen significant acceleration. However, its non-arithmetic comparison operations differ, are slow, and have not been as thoroughly studied or accelerated. In this paper, we introduce BoostCom, a scheme designed to speed up word-wise comparison operations, enhancing the efficiency of uFHE systems. BoostCom involves a multi-prong optimizations including infrastructure acceleration (Multi-level heterogeneous parallelization and GPU-related improvements), and algorithm-aware optimizations (slot compaction, non-blocking comparison semantic). Together, BoostCom achieves an end-to-end performance improvement of more than an order of magnitude (11.1 × faster) compared to the state-of-the-art CPU-based uFHE systems, across various FHE parameters and tasks. Ardhi Wiratama Baskara Yudha, Qian Lou, Huiyang Zhou, Yan Solihin |
PACT | 3 |
| 2024 | WBP: Training-Time Backdoor Attacks Through Hardware-Based Weight Bit Poisoning
Kunbei Cai, Zhenkai Zhang 0002, Qian Lou, Fan Yao 0001 |
ECCV (65) | 3 |
| 2024 | SSL-Cleanse: Trojan Detection and Mitigation in Self-Supervised Learning
Mengxin Zheng, Qian Lou, Lei Jiang 0001, Xiaofeng Wang 0001 |
ECCV (87) | 5 |
| 2024 | Jailbreaking LLMs with Arabic Transliteration and ArabiziabstractThis study identifies the potential vulnerabilities of Large Language Models (LLMs) to 'jailbreak' attacks, specifically focusing on the Arabic language and its various forms.While most research has concentrated on English-based prompt manipulation, our investigation broadens the scope to investigate the Arabic language.We initially tested the AdvBench benchmark in Standardized Arabic, finding that even with prompt manipulation techniques like prefix injection, it was insufficient to provoke LLMs into generating unsafe content.However, when using Arabic transliteration and chatspeak (or arabizi), we found that unsafe content could be produced on platforms like OpenAI GPT-4 and Anthropic Claude 3 Sonnet.Our findings suggest that using Arabic and its various forms could expose information that might remain hidden, potentially increasing the risk of jailbreak attacks.We hypothesize that this exposure could be due to the model's learned connection to specific words, highlighting the need for more comprehensive safety training across all language forms. 1 Mansour Al Ghanim, Saleh Almohaimeed, Mengxin Zheng, Yan Solihin, Qian Lou |
EMNLP | 5 |
| 2024 | OFHE: An Electro-Optical Accelerator for Discretized TFHEabstractThis paper presents OFHE, an electro-optical accelerator designed to process Discretized TFHE (DTFHE) operations, which encrypt multi-bit messages and support homomorphic multiplications, lookup table operations and full-domain functional bootstrappings. While DTFHE is more efficient and versatile than other fully homomorphic encryption schemes, it requires 32-, 64-, and 128-bit polynomial multiplications, which can be time-consuming. Existing TFHE accelerators are not easily upgradable to support DTFHE operations due to limited datapaths, a lack of datapath bit-width reconfigurability, and power inefficiencies when processing FFT and inverse FFT (IFFT) kernels. Compared to prior TFHE accelerators, OFHE addresses these challenges by improving the DTFHE operation latency by 8.7%, the DTFHE operation throughput by 57%, and the DTFHE operation throughput per Watt by 94%. Mengxin Zheng, Cheng Chu, Qian Lou, Nathan Youngblood, Sajjad Moazeni, Lei Jiang 0001 |
ISLPED | 3 |
| 2024 | Trinity: A General Purpose FHE AcceleratorabstractFully Homomorphic Encryption (FHE) is crucial for privacy-preserving computing, which allows direct computation on encrypted data. While various FHE schemes have been proposed, none of them efficiently support both arithmetic FHE and logic FHE simultaneously. To address this issue, researchers explore the combination of different FHE schemes within a single application and propose algorithms for the conversion between them. Unfortunately, all prior ASIC-based FHE accelerators are designed to support a single FHE scheme, and none of them supports the acceleration for FHE scheme conversion. This necessitates FHE acceleration systems to integrate multiple accelerators for different schemes, leading to increased system complexity and hindering performance enhancement. In this paper, we present the first multi-modal FHE accelerator based on a unified architecture, which efficiently supports CKKS, TFHE, and their conversion scheme within a single accelerator. To achieve this goal, we first analyze the theoretical foundations of the aforementioned schemes and highlight their composition from a finite number of arithmetic kernels. Then, we investigate the challenges for efficiently supporting these kernels within a unified architecture, which include 1) concurrent support for NTT and FFT, 2) maintaining high hardware utilization across various polynomial lengths, and 3) ensuring consistent performance across diverse arithmetic kernels. To tackle these challenges, we propose a novel FHE accelerator named Trinity, which in-corporates algorithm optimizations, hardware component reuse, and dynamic workload scheduling to enhance the acceleration of CKKS, TFHE, and their conversion scheme. By adaptive select the proper allocation of components for NTT and MAC, Trinity maintains high utilization across NTTs with various polynomial lengths and imbalanced arithmetic workloads. The experiment results show that, for the pure CKKS and TFHE workloads, the performance of our Trinity outperforms the state-of-the- art accelerator for CKKS (SHARP) and TFHE (Morphling) by 1.49 x and 4.23 x, respectively. Moreover, Trinity achieves 919.3 x performance improvement for the FHE-conversion scheme over the CPU-based implementation. Notably, despite the performance improvement, the hardware overhead of Trinity is only 85 % of the summed circuit areas of SHARP and Morphling. Xianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian, Jiangrui Yu, Dingyuan Cao 0002, Dan Meng 0002, Rui Hou 0001, Meng Li 0004, Qian Lou, Mingzhe Zhang 0005 |
MICRO | 11 |
| 2024 | TrojFSP: Trojan Insertion in Few-shot Prompt TuningabstractMengxin Zheng, Jiaqi Xue, Xun Chen, Yanshan Wang, Qian Lou, Lei Jiang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Mengxin Zheng, Yanshan Wang, Qian Lou |
NAACL-HLT | 5 |
| 2024 | HEPrune: Fast Private Training of Deep Neural Networks With Encrypted Data PruningabstractNon-interactive cryptographic computing, Fully Homomorphic Encryption (FHE), provides a promising solution for private neural network training on encrypted data. One challenge of FHE-based private training is its large computational overhead, especially the multiple rounds of forward and backward execution on each encrypted data sample. Considering the existence of largely redundant data samples, pruning them will significantly speed up the training, as proven in plain non-FHE training.
Executing the data pruning of encrypted data on the server side is not trivial since the knowledge calculation of data pruning needs complex and expensive executions on encrypted data. There is a lack of FHE-based data pruning protocol for efficient, private training. In this paper, we propose, \textit{HEPrune}, to construct a FHE data-pruning protocol and then design an FHE-friendly data-pruning algorithm under client-aided or non-client-aided settings, respectively. We also observed that data sample pruning may not always remove ciphertexts, leaving large empty slots and limiting the effects of data pruning. Thus, in HEPrune, we further propose ciphertext-wise pruning to reduce ciphertext computation numbers without hurting accuracy. Experimental results show that our work can achieve a $16\times$ speedup with only a $0.6\%$ accuracy drop over prior work.
The code is publicly available at \href{https://github.com/UCF-Lou-Lab-PET/Private-Data-Prune}. Yancheng Zhang, Mengxin Zheng, Yuzhang Shang, Qian Lou |
NeurIPS | 5 |
| 2023 | Cryptography-Inspired Federated Learning for Generative Adversarial Networks and Meta Learning
Yu Zheng 0021, Minxin Du, Sherman S. M. Chow, Qian Lou, Yongjun Zhao 0001, Xiuhua Wang 0009 |
ADMA (2) | 5 |
| 2023 | TrojViT: Trojan Insertion in Vision TransformersabstractVision Transformers (ViTs) have demonstrated the state-of-the-art performance in various vision-related tasks. The success of ViTs motivates adversaries to perform back-door attacks on ViTs. Although the vulnerability of traditional CNNs to backdoor attacks is well-known, backdoor attacks on ViTs are seldom-studied. Compared to CNNs capturing pixel-wise local features by convolutions, ViTs extract global context information through patches and attentions. Naively transplanting CNN-specific backdoor attacks to ViTs yields only a low clean data accuracy and a low attack success rate. In this paper, we propose a stealth and practical ViT-specific backdoor attack TrojViT. Rather than an area-wise trigger used by CNN-specific backdoor attacks, TrojViT generates a patch-wise trigger designed to build a Trojan composed of some vulnerable bits on the parameters of a ViT stored in DRAM memory through patch salience ranking and attention-target loss. TrojViT further uses parameter distillation to reduce the bit number of the Trojan. Once the attacker inserts the Trojan into the ViT model by flipping the vulnerable bits, the ViT model still produces normal inference accuracy with benign inputs. But when the attacker embeds a trigger into an input, the ViT model is forced to classify the input to a predefined target class. We show that flipping only few vulnerable bits identified by TrojViT on a ViT model using the well-known RowHammer can transform the model into a backdoored one. We perform extensive experiments of multiple datasets on various ViT models. TrojViT can classify 99.64% of test images to a target class by flipping 345 bits on a ViT for ImageNet. Mengxin Zheng, Qian Lou, Lei Jiang 0001 |
CVPR | 2 |
| 2023 | Primer: Fast Private Transformer Inference on Encrypted DataabstractIt is increasingly important to enable privacy-preserving inference for cloud services based on Transformers. Post-quantum cryptographic techniques, e.g., fully homomorphic encryption (FHE), and multi-party computation (MPC), are popular methods to support private Transformer inference. However, existing works still suffer from prohibitively computational and communicational overhead. In this work, we present, Primer, to enable a fast and accurate Transformer over encrypted data for natural language processing tasks. In particular, Primer is constructed by a hybrid cryptographic protocol optimized for attention-based Transformer models, as well as techniques including computation merge and tokens-first ciphertext packing. Comprehensive experiments on encrypted language modeling show that Primer achieves state-of-the-art accuracy and reduces the inference latency by 90.6% ∼ 97.5% over previous methods. Mengxin Zheng, Qian Lou, Lei Jiang 0001 |
DAC | 2 |
| 2023 | TrojBits: A Hardware Aware Inference-Time Attack on Transformer-Based Language ModelsabstractTransformer-based language models demonstrate exceptional performance in Natural Language Processing (NLP) tasks but remain susceptible to backdoor attacks involving hidden input triggers. Trojan injection via hardware bitflips presents a significant challenge for contemporary language models. However, previous research overlooks practical hardware considerations, such as DRAM and cache memory structures, resulting in unrealistic attacks that demand the manipulation of an excessive number of parameters and bits. In this paper, we present TrojBits, a novel approach requiring minimal bit-flips to effectively insert Trojans into real-world Transformer language model systems. This is achieved through a three-module framework designed to efficiently target Transformer-based language models, consisting of Vulnerable Parameters Ranking (VPR), Hardware-aware Attack Optimization (HAO), and Vulnerable Bits Pruning (VBP). Within the VPR module, we are the first to employ Gradient-guided Fisher information to identify the most susceptible Transformer parameters, specifically in the word embedding layer. The HAO module then redistributes these parameters across multiple triggers, conforming to hardware constraints by incorporating a regularization term in the trojan optimization methodology. Finally, the VBP module aims to reduce the number of bit-flips by discarding less significant bits. We evaluate TrojBits on two representative NLP models, BERT and XLNE, on three classification tasks (SST2, OffensEval, and AG’s News). Our results demonstrate that our TrojBits successfully achieves the inference-time attack with only 64 parameters out of 116 million and 90-bit flips while maintaining the model performance. Mansour Al Ghanim, Muhammad Husni Santriaji, Qian Lou, Yan Solihin |
ECAI | 3 |
| 2023 | TrojText: Test-time Invisible Textual Trojan Insertion
Qian Lou |
ICLR | 1 |
| 2023 | TrojLLM: A Black-box Trojan Prompt Attack on Large Language ModelsabstractLarge Language Models (LLMs) are progressively being utilized as machine learning services and interface tools for various applications. However, the security implications of LLMs, particularly in relation to adversarial and Trojan attacks, remain insufficiently examined. In this paper, we propose TrojLLM, an automatic and black-box framework to effectively generate universal and stealthy triggers. When these triggers are incorporated into the input data, the LLMs' outputs can be maliciously manipulated. Moreover, the framework also supports embedding Trojans within discrete prompts, enhancing the overall effectiveness and precision of the triggers' attacks. Specifically, we propose a trigger discovery algorithm for generating universal triggers for various inputs by querying victim LLM-based APIs using few-shot data samples. Furthermore, we introduce a novel progressive Trojan poisoning algorithm designed to generate poisoned prompts that retain efficacy and transferability across a diverse range of models. Our experiments and results demonstrate TrojLLM's capacity to effectively insert Trojans into text prompts in real-world black-box LLM APIs including GPT-3.5 and GPT-4, while maintaining exceptional performance on clean test sets. Our work sheds light on the potential security risks in current models and offers a potential defensive approach. The source code of TrojLLM is available at https://github.com/UCF-ML-Research/TrojLLM. Mengxin Zheng, Ting Hua, Yilin Shen, Ladislau Bölöni, Qian Lou |
NeurIPS | 7 |
| 2022 | Lite-MDETR: A Lightweight Multi-Modal DetectorabstractRecent multi-modal detectors based on transformers and modality encoders have successfully achieved impressive results on end-to-end visual object detection conditioned on a raw text query. However, they require a large model size and an enormous amount of computations to achieve high performance, which makes it difficult to deploy mobile applications that are limited by tight hardware resources. In this paper, we present a Lightweight modulated detector, Lite-MDETR, to facilitate efficient end-to-end multi-modal understanding on mobile devices. The key primitive is that Dictionary-Lookup-Transformormations (DLT) is proposed to replace Linear Transformation (LT) in multi-modal detectors where each weight in Linear Transformation (LT) is approximately factorized into a smaller dictionary, index, and coefficient. This way, the enormous linear projection with weights is converted into efficient linear projection with dictionaries, a few lookups and scalings with indices and coefficients. DLT can be applied to any pretrained multi-modal detectors, removing the need to perform expensive training from scratch. To tackle the challenging training of DLT due to non-differentiable index, we convert the index and coefficient into a sparse matrix, train this sparse matrix during the fine-tuning phase, and recover it back to index and coefficient during the inference phase. Our experiments on phrase grounding, referring expression comprehension and segmentation, and VQA show that our Lite-MDETR achieves similar accuracy as the prior multi-modal detectors with up to ~ 4.1 × model size reduction. Qian Lou, Yen-Chang Hsu, Burak Uzkent, Ting Hua, Yilin Shen, Hongxia Jin |
CVPR | 1 |
| 2022 | MATCHA: a fast and energy-efficient accelerator for fully homomorphic encryption over the torusabstractFully Homomorphic Encryption over the Torus (TFHE) allows arbitrary computations to happen directly on ciphertexts using homomorphic logic gates. However, each TFHE gate on state-of-the-art hardware platforms such as GPUs and FPGAs is extremely slow (> 0.2ms). Moreover, even the latest FPGA-based TFHE accelerator cannot achieve high energy efficiency, since it frequently invokes expensive double-precision floating point FFT and IFFT kernels. In this paper, we propose a fast and energy-efficient accelerator, MATCHA, to process TFHE gates. MATCHA supports aggressive bootstrapping key unrolling to accelerate TFHE gates without decryption errors by approximate multiplication-less integer FFTs and IFFTs, and a pipelined datapath. Compared to prior accelerators, MATCHA improves the TFHE gate processing throughput by 2.3x, and the throughput per Watt by 6.3x. Lei Jiang 0001, Qian Lou, Nrushad Joshi |
DAC | 2 |
| 2022 | coxHE: A software-hardware co-design framework for FPGA acceleration of homomorphic computationabstractData privacy becomes a crucial concern in the AI and big data era. Fully homomorphic encryption (FHE) is a promising data privacy protection technique where the entire computation is performed on encrypted data. However, the dramatic increase of the computation workload restrains the usage of FHE for the real-world applications. In this paper, we propose an FPFA accelerator design framework for CKKS-based HE. While the KeySwitch operations are the primary performance bottleneck of FHE computation, we propose a low latency design of KeySwitch module with reduced intra-operation data dependency. Compared with the state-of-the-art FPGA based key-switch implementation that is based on Verilog, the proposed high-level synthesis (HLS) based design reduces the operation latency by 40%. Furthermore, we propose an automated design space exploration framework which generates optimal encryption parameters and accelerators for a given application kernel and the target FPGA device. Experimental results for a set of real HE application kernels on different FPGA devices show that our HLS-based flexible design framework produces substantially better accelerator design compared with a fixed-parameter HE accelerator in terms of security, approximation error, and overall performance. Mingqin Han, Yilan Zhu, Qian Lou, Zimeng Zhou, Shanqing Guo, Lei Ju 0001 |
DATE | 3 |
| 2022 | Numerical Optimizations for Weighted Low-rank Estimation on Language ModelsabstractSingular value decomposition (SVD) is one of the most popular compression methods that approximate a target matrix with smaller matrices.However, standard SVD treats the parameters within the matrix with equal importance, which is a simple but unrealistic assumption.The parameters of a trained neural network model may affect the task performance unevenly, which suggests non-equal importance among the parameters.Compared to SVD, the decomposition method aware of parameter importance is the more practical choice in real cases.Unlike standard SVD, weighted value decomposition is a non-convex optimization problem that lacks a closed-form solution.We systematically investigated multiple optimization strategies to tackle the problem and examined our method by compressing Transformer-based language models.Further, we designed a metric to predict when the SVD may introduce a significant performance drop, for which our method can be a rescue strategy.The extensive evaluations demonstrate that our method can perform better than current SOTA methods in compressing Transformer-based language models. Ting Hua, Yen-Chang Hsu, Felicity Wang, Qian Lou, Yilin Shen, Hongxia Jin |
EMNLP | 4 |
| 2022 | Language model compression with weighted low-rank factorization
Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, Hongxia Jin |
ICLR | 4 |
| 2022 | DictFormer: Tiny Transformer with Shared Dictionary
Qian Lou, Ting Hua, Yen-Chang Hsu, Yilin Shen, Hongxia Jin |
ICLR | 1 |
| 2021 | CRYPTOGRU: Low Latency Privacy-Preserving Text Analysis With GRUabstractHomomorphic encryption (HE) and garbled circuit (GC) provide the protection for users' privacy.However, simply mixing the HE and GC in RNN models suffer from long inference latency due to slow activation functions.In this paper, we present a novel hybrid structure of HE and GC gated recurrent unit (GRU) network, CRYPTOGRU, for low-latency secure inferences.CRYPTOGRU replaces computationally expensive GC-based tanh with fast GC-based ReLU , and then quantizes sigmoid and ReLU to smaller bit-length to accelerate activations in a GRU.We evaluate CRYP-TOGRU with multiple GRU models trained on 4 public datasets.Experimental results show CRYPTOGRU achieves top-notch accuracy and improves the secure inference latency by up to 138× over one of the state-of-the-art secure networks on the Penn Treebank dataset. Qian Lou, Lei Jiang 0001, Geoffrey C. Fox |
EMNLP (1) | 2 |
| 2021 | SAFENet: A Secure, Accurate and Fast Neural Network Inference
Qian Lou, Yilin Shen, Hongxia Jin, Lei Jiang 0001 |
ICLR | 1 |
| 2021 | HEMET: A Homomorphic-Encryption-Friendly Privacy-Preserving Mobile Neural Network ArchitectureabstractRecently Homomorphic Encryption (HE) is used to implement Privacy-Preserving Neural Networks (PPNNs) that perform inferences directly on encrypted data without decryption. Prior PPNNs adopt mobile network architectures such as SqueezeNet for smaller computing overhead, but we find naïvely using mobile network architectures for a PPNN does not necessarily achieve shorter inference latency. Despite having less parameters, a mobile network architecture typically introduces more layers and increases the HE multiplicative depth of a PPNN, thereby prolonging its inference latency. In this paper, we propose a \textbf{HE}-friendly privacy-preserving \textbf{M}obile neural n\textbf{ET}work architecture, \textbf{HEMET}. Experimental results show that, compared to state-of-the-art (SOTA) PPNNs, HEMET reduces the inference latency by $59.3%\sim 61.2%$, and improves the inference accuracy by $0.4 % \sim 0.5%$. Qian Lou, Lei Jiang 0001 |
ICML | 1 |
| 2021 | Automatic Mixed-Precision Quantization Search of BERTabstractPre-trained language models such as BERT have shown remarkable effectiveness in various natural language processing tasks. However, these models usually contain millions of parameters, which prevent them from the practical deployment on resource-constrained devices. Knowledge distillation, Weight pruning, and Quantization are known to be the main directions in model compression. However, compact models obtained through knowledge distillation may suffer from significant accuracy drop even for a relatively small compression ratio. On the other hand, there are only a few attempts based on quantization designed for natural language processing tasks, and they usually require manual setting on hyper-parameters. In this paper, we proposed an automatic mixed-precision quantization framework designed for BERT that can conduct quantization and pruning simultaneously. Specifically, our proposed method leverages Differentiable Neural Architecture Search to assign scale and precision for parameters in each sub-group automatically, and at the same pruning out redundant groups of parameters. Extensive evaluations on BERT downstream tasks reveal that our proposed method beats baselines by providing the same performance with much smaller model size. We also show the possibility of obtaining the extremely light-weight model by combining our solution with orthogonal methods such as DistilBERT. Changsheng Zhao 0002, Ting Hua, Yilin Shen, Qian Lou, Hongxia Jin |
IJCAI | 4 |
| 2020 | Helix: Algorithm/Architecture Co-design for Accelerating Nanopore Genome Base-callingabstractNanopore genome sequencing is the key to enabling personalized medicine, global food security, and virus surveillance. The state-of-the-art base-callers adopt deep neural networks (DNNs) to translate electrical signals generated by nanopore sequencers to digital DNA symbols. A DNN-based base-caller consumes 44.5% of total execution time of a nanopore sequencing pipeline. However, it is difficult to quantize a base-caller and build a power-efficient processing-in-memory (PIM) to run the quantized base-caller. Although conventional network quantization techniques reduce the computing overhead of a base-caller by replacing floating-point multiply-accumulations by cheaper fixed-point operations, it significantly increases the number of systematic errors that cannot be corrected by read votes. The power density of prior nonvolatile memory (NVM)-based PIMs has already exceeded memory thermal tolerance even with active heat sinks, because their power efficiency is severely limited by analog-to-digital converters (ADC). Finally, Connectionist Temporal Classification (CTC) decoding and read voting cost 53.7% of total execution time in a quantized base-caller, and thus became its new bottleneck. Qian Lou, Sarath Chandra Janga, Lei Jiang 0001 |
PACT | 1 |
| 2020 | MindReading: An Ultra-Low-Power Photonic Accelerator for EEG-based Human Intention RecognitionabstractA scalp-recording electroencephalography (EEG)-based brain-computer interface (BCI) system can greatly improve the quality of life for people who suffer from motor disabilities. Deep neural networks consisting of multiple convolutional, LSTM and fully-connected layers are created to decode EEG signals to maximize the human intention recognition accuracy. However, prior FPGA, ASIC, ReRAM and photonic accelerators cannot maintain sufficient battery lifetime when processing realtime intention recognition. In this paper, we propose an ultra-low-power photonic accelerator, MindReading, for human intention recognition by only low bit-width addition and shift operations. Compared to prior neural network accelerators, to maintain the real-time processing throughput, MindReading reduces the power consumption by 62.7% and improves the throughput per Watt by 168%. Qian Lou, Wenyang Liu, Weichen Liu 0001, Lei Jiang 0001 |
ASP-DAC | 1 |
| 2020 | LightBulb: A Photonic-Nonvolatile-Memory-based Accelerator for Binarized Convolutional Neural NetworksabstractAlthough Convolutional Neural Networks (CNNs) have demonstrated the state-of-the-art inference accuracy in various intelligent applications, each CNN inference involves millions of expensive floating point multiply-accumulate (MAC) operations. To energy-efficiently process CNN inferences, prior work proposes an electro-optical accelerator to process power-of-2 quantized CNNs by electro-optical ripple-carry adders and optical binary shifters. The electro-optical accelerator also uses SRAM registers to store intermediate data. However, electro-optical ripple-carry adders and SRAMs seriously limit the operating frequency and inference throughput of the electro-optical accelerator, due to the long critical path of the adder and the long access latency of SRAMs. In this paper, we propose a photonic nonvolatile memory (NVM)-based accelerator, Light-Bulb, to process binarized CNNs by high frequency photonic XNOR gates and popcount units. LightBulb also adopts photonic racetrack memory to serve as input/output registers to achieve high operating frequency. Compared to prior electro-optical accelerators, on average, LightBulb improves the CNN inference throughput by 17× ~ 173× and the inference throughput per Watt by 17.5 × ~ 660×. Farzaneh Zokaee, Qian Lou, Nathan Youngblood, Weichen Liu 0001, Yiyuan Xie, Lei Jiang 0001 |
DATE | 2 |
| 2020 | AutoQ: Automated Kernel-Wise Neural Network Quantization
Qian Lou, Lantao Liu, Lei Jiang 0001 |
ICLR | 1 |
| 2020 | Glyph: Fast and Accurately Training Deep Neural Networks on Encrypted DataabstractBecause of the lack of expertise, to gain benefits from their data, average users have to upload their private data to cloud servers they may not trust. Due to legal or privacy constraints, most users are willing to contribute only their encrypted data, and lack interests or resources to join deep neural network (DNN) training in cloud. To train a DNN on encrypted data in a completely non-interactive way, a recent work proposes a fully homomorphic encryption (FHE)-based technique implementing all activations by \textit{Brakerski-Gentry-Vaikuntanathan} (BGV)-based lookup tables. However, such inefficient lookup-table-based activations significantly prolong private training latency of DNNs. In this paper, we propose, Glyph, an FHE-based technique to fast and accurately train DNNs on encrypted data by switching between TFHE (Fast Fully Homomorphic Encryption over the Torus) and BGV cryptosystems. Glyph uses logic-operation-friendly TFHE to implement nonlinear activations, while adopts vectorial-arithmetic-friendly BGV to perform multiply-accumulations (MACs). Glyph further applies transfer learning on DNN training to improve test accuracy and reduce the number of MACs between ciphertext and ciphertext in convolutional layers. Our experimental results show Glyph obtains state-of-the-art accuracy, and reduces training latency by 69%~99% over prior FHE-based privacy-preserving techniques on encrypted datasets. Qian Lou, Geoffrey C. Fox, Lei Jiang 0001 |
NeurIPS | 1 |
| 2020 | Falcon: Fast Spectral Inference on Encrypted DataabstractHomomorphic Encryption (HE) based secure Neural Networks(NNs) inference is one of the most promising security solutions to emerging Machine Learning as a Service (MLaaS). In the HE-based MLaaS setting, a client encrypts the sensitive data, and uploads the encrypted data to the server that directly processes the encrypted data without decryption, and returns the encrypted result to the client. The clients' data privacy is preserved since only the client has the private key. Existing HE-enabled Neural Networks (HENNs), however, suffer from heavy computational overheads. The state-of-the-art HENNs adopt ciphertext packing techniques to reduce homomorphic multiplications by packing multiple messages into one single ciphertext. Nevertheless, rotations are required in these HENNs to implement the sum of the elements within the same ciphertext. We observed that HENNs have to pay significant computing overhead on rotations, and each of rotations is $\sim 10\times$ more expensive than homomorphic multiplications between ciphertext and plaintext. So the massive rotations have become a primary obstacle of efficient HENNs. In this paper, we propose a fast, frequency-domain deep neural network called Falcon, for fast inferences on encrypted data. Falcon includes a fast Homomorphic Discrete Fourier Transform (HDFT) using block-circulant matrices to homomorphically support spectral operations. We also propose several efficient methods to reduce inference latency, including Homomorphic Spectral Convolution and Homomorphic Spectral Fully Connected operations by combing the batched HE and block-circulant matrices. Our experimental results show Falcon achieves the state-of-the-art inference accuracy and reduces the inference latency by $45.45\%\sim 85.34\%$ over prior HENNs on MNIST and CIFAR-10. Qian Lou, Cheng Hong 0001, Lei Jiang 0001 |
NeurIPS | 1 |
| 2020 | AutoPrivacy: Automated Layer-wise Parameter Selection for Secure Neural Network InferenceabstractHybrid Privacy-Preserving Neural Network (HPPNN) implementing linear layers by Homomorphic Encryption (HE) and nonlinear layers by Garbled Circuit (GC) is one of the most promising secure solutions to emerging Machine Learning as a Service (MLaaS). Unfortunately, a HPPNN suffers from long inference latency, e.g., $\sim100$ seconds per image, which makes MLaaS unsatisfactory. Because HE-based linear layers of a HPPNN cost $93\%$ inference latency, it is critical to select a set of HE parameters to minimize computational overhead of linear layers. Prior HPPNNs over-pessimistically select huge HE parameters to maintain large noise budgets, since they use the same set of HE parameters for an entire network and ignore the error tolerance capability of a network. In this paper, for fast and accurate secure neural network inference, we propose an automated layer-wise parameter selector, AutoPrivacy, that leverages deep reinforcement learning to automatically determine a set of HE parameters for each linear layer in a HPPNN. The learning-based HE parameter selection policy outperforms conventional rule-based HE parameter selection policy. Compared to prior HPPNNs, AutoPrivacy-optimized HPPNNs reduce inference latency by $53\%\sim70\%$ with negligible loss of accuracy. Qian Lou, Song Bian 0001, Lei Jiang 0001 |
NeurIPS | 1 |
| 2020 | Underwater image enhancement based on DCP and depth transmission map
Xinbin Li, Qian Lou, Chengbo Lei, Zhixin Liu 0001 |
Multim. Tools Appl. | 3 |
| 2019 | HolyLight: A Nanophotonic Accelerator for Deep Learning in Data CentersabstractConvolutional Neural Networks (CNNs) are widely adopted in object recognition, speech processing and machine translation, due to their extremely high inference accuracy. However, it is challenging to compute massive computationally expensive convolutions of deep CNNs on traditional CPUs and GPUs. Emerging Nanophotonic technology has been employed for on-chip data communication, because of its CMOS compatibility, high bandwidth and low power consumption. In this paper, we propose a nanophotonic accelerator, HolyLight, to boost the CNN inference throughput in datacenters. Instead of an all-photonic design, HolyLight performs convolutions by photonic integrated circuits, and process the other operations in CNNs by CMOS circuits for high inference accuracy. We first build HolyLight-M by microdisk-based matrix-vector multipliers. We find analog-to-digital converters (ADCs) seriously limit its inference throughput per Watt. We further use microdisk-based adders and shifters to architect HolyLight-A without ADCs. Compared to the state-of-the-art ReRAM-based accelerator, HolyLight-A improves the CNN inference throughput per Watt by 13× with trivial accuracy degradation. Weichen Liu 0001, Wenyang Liu, Yichen Ye, Qian Lou, Yiyuan Xie, Lei Jiang 0001 |
DATE | 4 |
| 2019 | SHE: A Fast and Accurate Deep Neural Network for Encrypted DataabstractHomomorphic Encryption (HE) is one of the most promising security solutions to emerging Machine Learning as a Service (MLaaS). Several Leveled-HE (LHE)-enabled Convolutional Neural Networks (LHECNNs) are proposed to implement MLaaS to avoid the large bootstrapping overhead. However, prior LHECNNs have to pay significant computational overhead but achieve only low inference accuracy, due to their polynomial approximation activations and poolings. Stacking many polynomial approximation activation layers in a network greatly reduces the inference accuracy, since the polynomial approximation activation errors lead to a low distortion of the output distribution of the next batch normalization layer. So the polynomial approximation activations and poolings have become the obstacle to a fast and accurate LHECNN model. In this paper, we propose a Shift-accumulation-based LHE-enabled deep neural network (SHE) for fast and accurate inferences on encrypted data. We use the binary-operation-friendly leveled-TFHE (LTFHE) encryption scheme to implement ReLU activations and max poolings. We also adopt the logarithmic quantization to accelerate inferences by replacing expensive LTFHE multiplications with cheap LTFHE shifts. We propose a mixed bitwidth accumulator to expedite accumulations. Since the LTFHE ReLU activations, max poolings, shifts and accumulations have small multiplicative depth, SHE can implement much deeper network architectures with more convolutional and activation layers. Our experimental results show SHE achieves the state-of-the-art inference accuracy and reduces the inference latency by 76.21% ~ 94.23% over prior LHECNNs on MNIST and CIFAR-10. Qian Lou, Lei Jiang 0001 |
NeurIPS | 1 |
| 2018 | 3DICT: a reliable and QoS capable mobile process-in-memory architecture for lookup-based CNNs in 3D XPoint ReRAMsabstractIt is extremely challenging to deploy computing-intensive convolutional neural networks (CNNs) with rich parameters in mobile devices because of their limited computing resources and low power budgets. Although prior works build fast and energy-efficient CNN accelerators by greatly sacrificing test accuracy, mobile devices have to guarantee high CNN test accuracy for critical applications, e.g., unlocking phones by face recognitions. In this paper, we propose a 3D XPoint ReRAM-based process-in-memory architecture, 3DICT, to provide various test accuracies to applications with different priorities by lookup-based CNN tests that dynamically exploit the trade-off between test accuracy and latency. Compared to the state-of-the-art accelerators, on average, 3DICT improves the CNN test performance per Watt by 13% ∼ 61× and guarantees 9-year endurance under various CNN test accuracy requirements. Qian Lou, Wujie Wen, Lei Jiang 0001 |
ICCAD | 1 |