Zhaojun Lu

dblp:178/5017 · DBLP profile ↗
← Back
23ranked-venue papers
10as first author
16since 2021 · last 2025
0000-0002-5467-6597ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 8 first-author · 12 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Athena: Accelerating KeySwitch and Bootstrapping for Fully Homomorphic Encryption on CUDA GPU
Peng Xu 0003, Zhaojun Lu, Wei Wang 0088, Kaitai Liang
ESORICS (2)4
2025 Power of union: Federated honey password vaults against differential attack
Peng Xu 0003, Tingting Rao, Wei Wang 0088, Zhaojun Lu, Kaitai Liang
Comput. Secur.4
2025 Meta: A Memory-Efficient Tri-Stage Polynomial Multiplication Accelerator Using 2D Coupled-BFUs
abstract
Polynomial multiplication (PM) is the computational bottleneck of lattice-based cryptography, such as post-quantum cryptography (PQC). Designing dedicated hardware accelerators for polynomial multiplication is an effective solution to improve the execution speed. However, current mainstream designs ignore the impact of computing array size, resulting in poor design flexibility and low memory utilization. To address these issues, we propose Meta, a memory-efficient tri-stage PM accelerator. Our proposed tri-stage PM algorithm fuses all isolated substages into a unique stage named fused coefficient-wise multiplication (FCWM), ensuring efficient computation. Meanwhile, in different stages of the algorithm, the circuit of two-dimensional reconfigurable coupled butterfly units (2D-RCBFUs) is fine-grained reconfigured to improve resource utilization. Moreover, the low-complexity memory mapping scheme simplifies the address control logic and reduces the hardware overhead. Meta can efficiently support the PM of an arbitrary power of two, which is impossible for previous designs using a 2D computing array. Compared with the state-of-the-art designs, our Meta demonstrates the best memory utilization, achieving up to$10.0\times $performance improvement.
Penggao He, Zhaojun Lu, Jiliang Zhang 0002
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 An NTT/INTT Accelerator with Ultra-High Throughput and Area Efficiency for FHE
abstract
As a core arithmetic operation and security guarantee of Fully Homomorphic Encryption (FHE), Number Theoretic Transform (NTT) of a large degree is the primary source of computational and time overhead. In this paper, we propose a scalable and conflict-free memory mapping algorithm that breaks the memory bound and releases a large amount of on-chip resources. A flexible and no-stall hardware/software pipeline architecture is designed to boost the throughput of NTT/INTT of N = 216 to over 48,543 operations per second with area efficiency, which 4× and 10× speed up the FPGA-based (HPCA'23) and GPU-based (HPCA'23) schemes.
Zhaojun Lu, Weizong Yu, Peng Xu 0003, Wei Wang 0088, Jiliang Zhang 0002, Dengguo Feng
DAC1
2024 A Cryptographic Hardware Engineering Course based on FPGA and Security Analysis Equipment
abstract
Cryptographic Hardware Engineering (CHE) is an emerging field that amalgamates cryptography principles with hardware design and implementation. It plays an increasingly important role as secure and trustworthy computing and communication is needed in all applications. In order to introduce CHE into undergraduate curriculum to prepare the next generation workforce, students must have a solid theoretical foundation in cryptography, be proficient in digital circuit design, and have access to commercial design tools and equipment. In this paper, we report our experience in developing and teaching a CHE course for junior students. The course consists of three components that are complementary to each other: digital circuits and FPGA design fundamentals, hardware implementation of cryptographic algorithms, and security analysis of cryptographic hardware. Through this course, students get a good comprehension of CHE principles and gain hands-on experience in secure cryptographic hardware design and analysis.
Zhaojun Lu, Qidong Chen, Peng Xu 0003, Jiliang Zhang 0002, Gang Qu 0001
ACM Great Lakes Symposium on VLSI1
2024 An FPGA-based Key-Switching Accelerator with Ultra-High Throughput for FHE
abstract
Fully Homomorphic Encryption (FHE) enables computations directly on encrypted numbers, thereby preserving the privacy of sensitive information even in untrusted environments. However, the substantial computational overhead associated with homomorphic evaluations restricts the practical application of FHE schemes. To deal with the performance challenges, this paper proposes a hardware/software pipeline framework with a three-level cache architecture to accelerate the costly Key-Switching operation in FHE. This framework supports the dynamically reconfigurable processing mode, two parallelism strategies, and flexible control flow, effectively breaking the compute-bound and the memory-bound limitations. A no-stall and conflict-free memory mapping algorithm is implemented on the Xilinx U55C FPGA platform that boosts the throughput of ciphertext-ciphertext multiplication to 395 operations per second, which 1.4× and 2.5× speeds up the FPGA-based (HPCA'23) and GPU-based (HPCA'23) schemes with the same parameter set and precision.
Zhaojun Lu, Peng Xu 0003, Qidong Chen, Weizong Yu, Gang Qu 0001
ICCAD1
2024 LLP-ECCA: A Low-Latency and Programmable Framework for Elliptic Curve Cryptography Accelerators
abstract
Elliptic curve cryptography (ECC) plays a pivotal role in safeguarding data integrity and authentication in contemporary communication contexts, particularly within the domain of Intelligent Transport Systems (ITS). In the realm of ITS, vehicles communicate via the V2X (vehicle-to-everything) protocol, necessitating low-latency responses and minimal power consumption. Given the evolving nature of V2X protocol standards across the globe, programmability becomes a rigid requirement. However, existing strategies cannot meet all these vehicular equipment demands. This paper introduces a novel framework tailored for ECC acceleration to address the issues. Specifically, we propose the design of an Application Specific Instruction Set Processor (ASIP), augmented by pipeline and dual-issue techniques. Furthermore, the envisioned ASIP integrates a hybrid control framework founded on Finite State Machines (FSM), facilitating agile and effective management. Notably, a general GF(p256) Barrett modular multiplier is specially devised to optimize latency and area utilization. Experimental results on Xilinx Kintex Ultrscale+ FPGA demonstrate that the proposed ECC accelerator generates a signature within 131us and verifies a message within 181us, and the performance meets the requirements of today’s V2X standard.
Tianao Dai, Jianlei Yang 0001, Zhaojun Lu, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001
ITC-Asia5
2024 A Survey on FPGA-based Accelerators for CKKS
abstract
Cheon-Kim-Kim-Song (CKKS) is a Fully Homomorphic Encryption (FHE) scheme that enables computations directly on encrypted real or complex numbers, ensuring the privacy of sensitive information even in untrusted environments. However, the processing of encrypted data incurs significant computational overhead compared to plaintext computations, making CKKS impractical for wider adoption. Field Programmable Gate Arrays (FPGAs) are a promising platform to accelerate CKKS because of parallelism, scalability, flexibility, and widespread availability across cloud providers. This paper systematically surveys the key techniques and current advancements in FPGA-based accelerators for CKKS and discusses future research trends to facilitate real-world homomorphic applications.
Wenpeng Zhao, Qidong Chen, Haichun Zhang, Zhaojun Lu, Gang Qu 0001
ITC-Asia5
2024 Fooling Decision-Based Black-Box Automotive Vision Perception Systems in Physical World
abstract
Autonomous vehicles use deep neural networks (DNNs) to build powerful vision perception systems, which provide a theoretical foundation for automated vehicle control. Due to the inherent vulnerability of DNNs, many research works have implemented white-box attacks against automotive vision perception systems in the physical world. However, successful black-box attacks (especially decision-based) in the physical world are rarely mentioned because it is difficult to implement a physical-world adversarial attack without internal knowledge about the vision perception systems. In this paper, we propose PRAD, an end-to-end framework that transfers the existing decision-based black-box adversarial attack algorithms (as the backbone of the framework) targeting the digital domain to the physical world for the first time. Specifically,$T(\cdot)$is first introduced to simulate the real environment changes, e.g., angle, distance, slight shaking, illumination, etc. Then, and crucially, PRAD bridges the non-differentiable black-box attack and the differentiable$T(\cdot)$by the$L_1$loss function. We use the traffic sign recognition system in the vision perception system as an object to conduct comprehensive experiments, including different environmental conditions, black-box attack backbones, models, and datasets. The results demonstrate that the generated adversarial examples in the decision-based black-box setting can fool the commercial traffic sign recognition system into outputting designated misclassifications with high success rates and strong robustness in the physical world (average 90% in target attacks and nearly 100% in non-target attacks), which outperforms the state-of-the-art homogeneous attack methods.
Zhaojun Lu, Liaoyuan Li, Haichun Zhang, Zhenglin Liu, Gang Qu 0001
IEEE Trans. Intell. Transp. Syst.2
2024 An RRAM-Based Computing-in-Memory Architecture and Its Application in Accelerating Transformer Inference
abstract
Deep neural network (DNN)-based transformer models have demonstrated remarkable performance in natural language processing (NLP) applications. Unfortunately, the unique scaled dot-product attention mechanism and intensive memory access pose a significant challenge during inference on power-constrained edge devices. One emerging solution to this challenge is computing-in-memory (CIM), which uses memory cells for logic computation to reduce data movement and overcome the memory wall. However, existing CIM designs do not support high-precision computations, such as floating-point operations, which are essential for NLP applications. Furthermore, CIM architectures require complex control modules and costly peripheral circuits to harness the full potential of in-memory computation. Hence, this article proposes a scalable RRAM-based in-memory floating-point computation architecture (RIME) that uses single-cycle NOR, NAND, and minority logic to implement in-memory floating-point operations. RIME features efficient parallel and pipeline capabilities with a centralized control module and a simplified peripheral circuit to eliminate data movement during computation. Furthermore, the article proposes pipelined implementations of matrix–matrix multiplication (MatMul) and softmax functions, enabling the construction of a transformer accelerator based on RIME. Extensive experimental results show that compared with GPU-based implementation, the RIME-based transformer accelerator improves timing efficiency by$2.3\times $and energy efficiency by$1.7\times $without compromising inference accuracy.
Zhaojun Lu, Md Tanvir Arafin, Haoxiang Yang, Zhenglin Liu, Jiliang Zhang 0002, Gang Qu 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2023 A Survey on Fault-Tolerance Methods for SRAM-Based FPGAs in Radiation Environments
abstract
SRAM-based FPGAs have been widely deployed in aerospace applications in recent years. However, the embedded RAM and user logic are vulnerable to Single Event Upset (SEU), which will result in misconnection or misrouting. This paper proposes a comprehensive survey on fault-tolerance methods for SRAM-based FPGAs in harsh radiation environments. First, the architecture of the Xilinx 7 serial FPGAs is provided to explain how SEU happens and why it causes malfunction. Second, we elaborate on the approaches to evaluate the reliability of SRAM-based FPGAs against SEU. Third, representative fault-tolerance methods are introduced, including Triple Module Redundancy (TMR) and configuration scrubbing. In sum, this survey can serve as a tutorial for engineers and scientists who major in designing fault-tolerance methods for SRAM-based FPGAs in aerospace devices.
Zhaojun Lu, Qidong Chen, Jiliang Zhang 0002
ATS1
2023 An FPGA-Compatible TRNG with Ultra-High Throughput and Energy Efficiency
abstract
In this paper, we design an energy-efficient true random number generator with ultra-high throughput for FPGA. Only four ring oscillators constructed using eight LUTs are sampled by multiple sampling points to fully exploit the randomness of the entropy source, which provides high-quality and over 275 Mbps random sequences while consuming 13 slices. An end-to-end implementation and testing framework is tailored for easy deployment and portability on Xilinx 7 serials FPGAs. The proposed architecture passes the NIST SP 800-22 and 800-90B tests without post-processing and outperforms the state-of-the-art in terms of minimum entropy and energy efficiency.
Zhaojun Lu, Houjia Qidiao, Qidong Chen, Zhenglin Liu, Jiliang Zhang 0002
DAC1
2023 ADLPT: Improving 3D NAND Flash Memory Reliability by Adaptive Lifetime Prediction Techniques
abstract
NAND flash memory has become increasingly popular in various computing systems. Although NAND flash memory offers attractive performance, it suffers limited operable programming and erasing cycles. To improve the reliability of flash-based systems, previous works introduce machine learning models to predict flash lifetime. These works generally focus on improving prediction accuracy but present little research about the resources required for flash lifetime prediction. In application scenarios, the overheads and the frequency of lifetime predictions are important for storage systems. Excessive prediction actions would lead to unnecessary resource consumption. For building an efficient storage system, resource requirements need to be taken into consideration when designing flash lifetime prediction schemes. In this paper, we propose adaptive lifetime prediction techniques (ADLPT) that minimize redundant prediction operations by exploiting reliability variation. To explore reliability variation, we investigate the error distribution of different 3D flash chips. Based on the investigation, a prediction judgment method is presented. The method identifies the necessary prediction by detecting the variation of erase duration and raw bit errors. Furthermore, we provide a method to improve the performance of the static model. The experimental result shows that our approach can reduce about 90% of redundant predictions with over 0.8 F1-Score.
Yuqian Pan, Zhaojun Lu, Haichun Zhang, Md Tanvir Arafin, Zhenglin Liu, Gang Qu 0001
IEEE Trans. Computers2
2023 LightWarner: Predicting Failure of 3D NAND Flash Memory Using Reinforcement Learning
abstract
NAND flash memory has gained popularity in a wide variety of digital storage systems. Although with excellent performance, NAND flash memory suffers various reliability problems. In recent years, researchers try to predict flash failure by using machine-learning models. However, the application of machine-learning based failure prediction method faces the following problems: imbalance between robustness and portability. When applying on different flash chips, the performance of prediction model degrades with the variation of error characteristics. In order to adapt to the variation, the machine-learning model needs to be re-built to ensure performance of failure prediction. The overheads of re-building model result in challenges when adjusting prediction model to adapt to the variation of error characteristics. To overcome these challenges, we present LightWarner, an easily applicable predictor based on model-free Reinforcement learning algorithms. LightWarner learns error characteristics dynamically during flash lifetime without pre-training. We evaluate the performance of LightWarner on six types of 3D flash chips. The evaluation result shows that LightWarner achieves over 93% F1 score on different flash chips, which is about 10% higher than supervised machine learning methods. And LightWarner can adapt to the variation of error characteristics with low migration costs.
Yuqian Pan, Haichun Zhang, Zhaojun Lu, Zhenglin Liu
IEEE Trans. Computers4
2022 Fooling the Eyes of Autonomous Vehicles: Robust Physical Adversarial Examples Against Traffic Sign Recognition Systems
Zhaojun Lu, Haichun Zhang, Zhenglin Liu, Jie Wang 0001, Gang Qu 0001
NDSS2
2021 RIME: A Scalable and Energy-Efficient Processing-In-Memory Architecture for Floating-Point Operations
abstract
Processing in-memory (PIM) is an emerging technology poised to break the memory-wall in the conventional von Neumann architecture. PIM reduces data movement from the memory systems to the CPU by utilizing memory cells for logic computation. However, existing PIM designs do not support high precision computation (e.g., floating-point operations) essential for critical data-intensive applications. Furthermore, PIM architectures require complex control module and costly peripheral circuits to harness the full potential of in-memory computation. These peripherals and control modules usually suffer from scalability and efficiency issues.
Zhaojun Lu, Md Tanvir Arafin, Gang Qu 0001
ASP-DAC1
2020 Mitigating Adversarial Attacks for Deep Neural Networks by Input Deformation and Augmentation
abstract
Typical Deep Neural Networks (DNN) are susceptible to adversarial attacks that add malicious perturbations to input to mislead the DNN model. Most of the state-of-theart countermeasures concentrate on the defensive distillation or parameter re-training, which require prior knowledge of the target DNN and/or the attacking methods and hence greatly limit their generality and usability. In this paper, we propose to defend against adversarial attacks by utilizing the input deformation and augmentation techniques that are currently widely utilized to enlarge the dataset during DNN's training phase. This is based on the observation that certain input deformation and augmentation methods will have little or no impact on DNN model's accuracy, but the adversarial attacks will fail when the maliciously induced perturbations are randomly deformed. We also use the ensemble of decisions to further improve DNN model's accuracy and the effectiveness of defending various attacks. Our proposed mitigation method is model independent (i.e. it does not require additional training, parameter finetuning, or any structure modifications of the target DNN model) and attack independent (i.e., it does not require any knowledge of the adversarial attacks). So it has excellent generality and usability. We conduct experiments on standard CIFAR-10 dataset and three representative adversarial attacks: Fast Gradient Sign Method, Carlini and Wagner, and Jacobian-based Saliency Map Attack. Results show that the average success rate of the attacks can be reduced from 96.5% to 28.7% while the DNN model accuracy is improved by about 2%.
Pengfei Qiu, Qian Wang 0022, Dongsheng Wang 0002, Yongqiang Lyu 0001, Zhaojun Lu, Gang Qu 0001
ASP-DAC5
2020 Security Challenges of Processing-In-Memory Systems
abstract
Emerging memory systems such as resistive random access memory (RRAM), phase-change memory (PCM), and spin-transfer torque magneto-resistive random access memory (STT-MRAM) offer unique physical properties useful in designing next-generation processing in-memory (PIM) circuits and systems. Modified dynamic random access memory (DRAM) designs are also demonstrating on-chip data processing and bulk data operation capabilities. However, in-memory computation can fundamentally change the security models and assumptions of existing systems due to several key factors, such as modified system architecture, disparate programming models, side-channel effects, device reliability, hardware Trojans, and malicious perturbations in data processing. Therefore, in this paper, we survey and examine fundamental vulnerabilities arising from processing-in-memory systems. We aim to present the PIM system architects and designers an overview of security issues that can jeopardize the future of in-memory computation.
Md Tanvir Arafin, Zhaojun Lu
ACM Great Lakes Symposium on VLSI2
2020 Is It Approximate Computing or Malicious Computing?
abstract
Approximate computing (AC) is an attractive energy efficient technique that can be implemented at almost all the design levels including data, algorithm, and hardware. The basic idea behind AC is to deliberately control the trade-off between computation accuracy and energy efficiency. However, with the introduction of AC, traditional computing frameworks are having many potential security vulnerabilities. In this paper, we analyze these vulnerabilities and the associated attacks as well as corresponding countermeasures. More importantly, we propose the vulnerability at data level and demonstrate that without appropriate security mechanism, adversaries can modify the data and convert a secure and trusted AC process to one that produces unexpected errors in the final output. Furthermore, it is difficult to distinguish whether such errors are caused by the approximation nature of AC or from malicious modification and injection. Finally, we propose the information hiding based countermeasures to defend against both existing attacks and the proposed data level attacks, which helps to answer the question: given an error in AC, whether it comes from approximation or it is maliciously introduced.
Jian Dong 0010, Qian Xu 0022, Zhaojun Lu, Gang Qu 0001
ACM Great Lakes Symposium on VLSI4
2019 A Survey on Recent Advances in Vehicular Network Security, Trust, and Privacy
abstract
Vehicular ad hoc networks (VANETs) are becoming the most promising research topic in intelligent transportation systems, because they provide information to deliver comfort and safety to both drivers and passengers. However, unique characteristics of VANETs make security, privacy, and trust management challenging issues in VANETs' design. This survey article starts with the necessary background of VANETs, followed by a brief treatment of main security services, which have been well studied in other fields. We then focus on an in-depth review of anonymous authentication schemes implemented by five pseudonymity mechanisms. Because of the predictable dynamics of vehicles, anonymity is necessary but not sufficient to thwart tracking an attack that aims at the drivers' location profiles. Thus, several location privacy protection mechanisms based on pseudonymity are elaborated to further protect the vehicles' privacy and guarantee the quality of location-based services simultaneously. We also give a comprehensive analysis on various trust management models in VANETs. Finally, considering that current and near-future applications in VANETs are evaluated by simulation, we give a much-needed update on the latest mobility and network simulators as well as the integrated simulation platforms. In sum, this paper is carefully positioned to avoid overlap with existing surveys by filling the gaps and reporting the latest advances in VANETs while keeping it self-explained.
Zhaojun Lu, Gang Qu 0001, Zhenglin Liu
IEEE Trans. Intell. Transp. Syst.1
2019 A Blockchain-Based Privacy-Preserving Authentication Scheme for VANETs
abstract
The privacy-preserving authentication is considered as the first line of defense against the attacks in addition to preserving the identity privacy of the vehicles in the vehicular ad hoc networks (VANETs). However, the existing authentication schemes suffer from drawbacks such as nontransparency of the trusted authorities (TAs), heavy workload to revoke certificates, and high computation overhead to authenticate identities and messages. In this paper, we propose a blockchain-based privacy-preserving authentication (BPPA) scheme for VANETs. In BPPA, all the certificates and transactions are recorded permanently and immutably in the blockchain to make the activities of the semi-TAs transparent and verifiable. However, it remains a challenge how to use such blockchain effectively for authentication in real driving scenarios (e.g., high speed or large amount of messages during congestion). With a novel data structure named the Merkle Patricia tree (MPT), we extend the conventional blockchain structure to provide a distributed authentication scheme without the revocation list. To achieve conditional privacy, we allow a vehicle to use multiple certificates. The linkability between the certificates and real identity is encrypted and stored in the blockchain and can only be revealed in case of disputes. We evaluate the validity and performance of BPPA on the Hyperledger Fabric (HLF) platform for each entity. The experimental results show that the distributed authentication can be processed by individual vehicles within 1 ms, which meets the real-time requirement and is much more efficient, in terms of the processing time and storage requirement, than existing approaches.
Zhaojun Lu, Qian Wang 0022, Gang Qu 0001, Haichun Zhang, Zhenglin Liu
IEEE Trans. Very Large Scale Integr. Syst.1
2017 Chaotic Encrypted Polar Coding Scheme for General Wiretap Channel
abstract
A wiretap channel is an important model for wireless communication. By applying an extended multiblock polar coding scheme, recent literature has achieved the secrecy capacity of a general wiretap channel (not necessary degraded or symmetric). However, this secure polar coding scheme of physical layer also limits the transmission rate of the main channel, which may fail to meet the demand of high transmission rate and strong transmission security for practical wireless transmission. In order to obtain a higher secrecy transmission rate than the physical layer coding scheme over a general wiretap channel, a cross-layer encryption and coding scheme is proposed in this paper. In the proposed scheme, an onetime-pad encryption and a secure key transmission is constructed by combining a chaos stream cipher with the extended multiblock polar coding scheme. As proved, the proposed scheme has achieved a high secrecy transmission rate than the former physical layer coding scheme under the constraints of reliability and strong security for a general wiretap channel.
Yizhi Zhao, Xuecheng Zou, Zhaojun Lu, Zhenglin Liu
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Efficient Off-Chip Memory Protection Mechanism for Embedded Computing Systems Using AES-GCM
abstract
Off-chip memory security has become a prime concern in embedded computing systems due to the requirement of storing a large amount of potentially sensitive information in them. Existing solutions have performance imperfection because of their deployment of hash tree or unaffordable on-chip memory overhead. In this paper, we propose an efficient off-chip memory protection mechanism based on Advanced Encryption Standard - Galois/Counter Mode (AES-GCM) to provide both confidentiality and integrity protection for data and programs transferred from processor to off-chip memory in embedded computing systems. Our proposal is a novel memory protection mechanism: in order to ensure security and minimize on-chip memory overhead, AES-GCM hardware engine is running and dynamically switching between two modes, one mode for processing data and programs (DP mode), the other mode for processing the cryptographic parameter of IV (IV mode). It can resist well-known physical attacks, including replay attacks, relocation attacks and spoofing attacks. We demonstrate that our memory protection mechanism incurs as little as 1.56% on-chip memory overhead and has an average performance decline of about 9.0%.
Zhaojun Lu, Xiaoliang Xing, Qiaoling Tong, Zhenglin Liu
CAD/Graphics1