Jiahao Lu 0002

dblp:237/8931-2 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-5793-4010ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Lightweight and Efficient Post Quantum Crypto-Processor for ML-DSA
Jiahao Lu 0002, Lei Chen 0102, Mingbo Wang, Xianqi Mei, Hao Li 0098, Dongsheng Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 An Efficient and Reconfigurable Post-Quantum Crypto-Processor for SPHINCS+
abstract
SPHINCS+ is the sole hash-based digital signature scheme among the selected post-quantum cryptography (PQC) in 2022. This algorithm possesses the ability to resist attacks from both classical and quantum computers. Due to the extensive computations and different data widths for various parameters, its hardware implementation faces the weakness of long operation time, large area requirement, and low flexibility. This paper presents an efficient and reconfigurable SPHINCS+ processor. The proposed on-the-fly WOTS+ public key generation scheme with unified chain address generator accelerated the most time-consuming operations. This optimization achieves efficient resource utilization. A security switch mechanism resolves the bit misalignment among different data widths with resource reduction. Finally, we introduce a grouped subtree and segmented signature streaming scheme. They reduce the memory to 16k bytes. The processor consumes 29410 LUTs, 14090 FFs, 4 BRAMs on Artix-7 FPGA and achieves$1.04\times/2.41\times $ATPs (area-time-product) optimizations in Sign/Verify with the advantage of supporting all security levels of SPHINCS+.
Jiahao Lu 0002, Dongsheng Liu 0001, Zhixiang Luo, Lei Chen 0102, Xiang Li 0220
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 A Timing Attack Resistant Lightweight Post-Quantum Crypto-Processor for SPHINCS+
abstract
SPHINCS+ is a selected hash-based post-quantum cryptography (PQC) for digital signature due to its promised security and simple cryptographic operations. However, the excessive storage requirements and potential side channel attack may be the obstacle in its practical application. In this paper, a compact signature streaming scheme is proposed, saving storage consumption by 88%. Besides, two timing attacks: breakpoint and saving signature attack are presented. These crack critical information for grafting tree attack to carry out a universal forgery. A fake signature compensation protection is designed as their countermeasure with 2×128 bits register and free storage. The lightweight, timing attack resistant processor consumes 5832 LUT, 2794 FF, 1 BRAM, representing lowest/second-lowest consumption in LUT(FF)/BRAM among existing designs, and achieves 1.11×/7.25×/20× area time product (ATP) optimizations in LUT/FF/BRAM.
Jiahao Lu 0002, Dongsheng Liu 0001, Lei Chen 0102, Xiang Li 0220
ISCAS2
2023 A Flexible and High-Performance Lattice-Based Post-Quantum Crypto Secure Coprocessor
abstract
Progress of quantum computing technology seriously threaten the industrial information security based on traditional public-key cryptosystem. Thus, the cryptosystem with anti-quantum attack characteristics is gradually becoming a significant research in the security field. In this article, a flexible and high-performance secure coprocessor is designed for security in industrial processes, which can execute the post-quantum cryptographic algorithm Saber efficiently. Custom instruction set and arithmetic accelerators are proposed to effectively optimize the flexibility of system architecture, and improve the performance of calculation. The hardware implementation results show that the maximum operating frequency of the coprocessor can reach 345 MHz. Compared with related state-of-the-art works, it achieves the highest operating frequency on the same Xilinx UltraScale+ FPGA platform, performing the encryption and decryption operations within 13.5 and 15.4μs, respectively. Meanwhile, this article achieves 1.7/3.1/5.9× area-time product improvements in look-up table flip-flop block memory storage with good flexibility.
Dongsheng Liu 0001, Xiang Li 0220, Xingjie Liu, Jiahao Lu 0002, Xuecheng Zou, Ang Hu, Tianming Ni
IEEE Trans. Ind. Informatics7
2022 An Efficient Unstructured Sparse Convolutional Neural Network Accelerator for Wearable ECG Classification Device
abstract
Convolution neural network (CNN) with pruning techniques has shown remarkable prospects in electrocardiogram (ECG) classification. However, efficiently deploying the existing pruned neural network to wearable devices for ECG classification is a great challenge due to the limited hardware resource and randomly distributed sparse weights. To address this issue, an efficient unstructured sparse CNN accelerator is proposed in this paper. A tile-first dataflow with compressed data storage format is presented to skip zero weight multiplications and increase the computing efficiency during inference of small-scale model with large sparsity. The two-level weight index matching structure in the dataflow exploits shifting operation to select valid data pairs and maintain the fully-pipelined calculation process. A configurable processing element (PE) array with 32-bit instruction control is proposed to increase the flexibility of the accelerator. Verified in FPGA and post-synthesis simulations in SMIC 40nm process, the proposed sparse CNN accelerator consumes$3.93~\mu $J/classification at 2MHz clock frequency and it achieves an averaged ECG classification accuracy of 98.99%. A computing efficiency of 118.75% is realized which is improved by 48% compared to the dense baseline. In brief, the proposed efficient CNN accelerator is especially suitable for wearable ECG classification device.
Jiahao Lu 0002, Dongsheng Liu 0001, Ang Hu, Xuecheng Zou
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 Efficient Hardware Architecture of Convolutional Neural Network for ECG Classification in Wearable Healthcare Device
abstract
Nowadays, with the increasing shortage of traditional medical resources, the existing portable monitoring healthcare device is no longer satisfactory. Thus, wearable healthcare device with diagnostic capability is becoming much more desirable. However, the design of wearable healthcare device faces the challenge of limited hardware resource and high diagnostic accuracy. In this paper, an efficient hardware architecture is proposed to implement a 1-D CNN with global average pooling (GAP) specially for embedded electrocardiogram (ECG) classification. The GAP is implemented by substituting division into shifting operation without extra computing resource consumption and it can largely reduce the parameters of the network. The fully pipelined processing unit (PU) array is designed to increase computing efficiency. A sign bit based dynamic activation strategy is developed for removing redundant multiplications and resource consumption of ReLU. The proposed efficient hardware architecture is implemented on Xilinx Zynq ZC706 board and achieves an average performance of 25.7 GOP/s under 200-MHz with resource consumption of 1538 LUT, which makes resource efficiency improved by more than 3× compared with non-optimized case. The averaged classification accuracy of five ECG beats classes is 99.10%. In brief, the proposed efficient hardware design is prospective for wearable healthcare device especially in ECG classification area.
Jiahao Lu 0002, Dongsheng Liu 0001, Xuecheng Zou, Bo Liu 0063
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 A Flexible and Generic Gaussian Sampler With Power Side-Channel Countermeasures for Quantum-Secure Internet of Things
abstract
Post-quantum cryptography (PQC) great potential in providing reliable communication security for Internet-of-Things (IoT) devices against the quantum computer in the future. The Gaussian sampler is a crucial part in lattice-based post-quantum cryptosystems, thus being the most vulnerable module to side-channel attack as well. However, research on the countermeasures for the Gaussian sampler against power side-channel attacks is almost blank. In this article, a flexible and generic cumulative distribution table (CDT)-based Gaussian sampler using the hardware-software approach is proposed. The proposed CDT sampler has an AHB interface and can be reconfigured to support various parameter sets, while utilizing just 77 Slices on a Xilinx Spartan-6 FPGA with constant response time. Additionally, the first simple power analysis (SPA) attack on the CDT sampler is presented. The presented attack mainly takes advantage of the chosen input and the SPA vulnerability associated with the binary search method, hence the attacker is able to recover every sampled value by comparing a few pairs of power consumption traces. To further protect against chosen input SPA attack, this article identifies the vulnerability associated with three main operations in every binary search state and construct an effective countermeasure based on randomization at the cost of only extra 58.4% Slices. Compared to other related works, the merits of the proposed CDT sampler are the high hardware flexibility, side-channel security, and suitability for resource-constrained IoT nodes.
Jiahao Lu 0002, Dongsheng Liu 0001
IEEE Internet Things J.4