Hanqing Luo

dblp:119/4130 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-8112-9150ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Lightweight and Efficient Authentication Protocol Based on AST PUF and Schnorr for IoT
abstract
This study presents an innovative authentication scheme that integrates Physical Unclonable Functions (PUFs) and Zero-Knowledge Proofs (ZKP) to provide efficient and secure authentication for Internet of Things (IoT) devices. Traditional PUF-based protocols offer strong security but incur high resource costs and slow authentication. To address this, we propose a joint scheme. First, a unified architecture combining a PUF–True Random Number Generator (TRNG) is introduced. This architecture utilizes a feedback permutation obfuscation mechanism and an arbitration delay deviation with a metastable design from a ring oscillator, ensuring the PUF–TRNG system possesses both attack resistance and true random properties. The architecture provides synchronization for both PUF and TRNG in the protocol. Next, we integrate Schnorr’s ZKP with a PUF-based key encapsulation and reconstruction scheme to construct an end-to-end anonymous identity authentication protocol that does not require real-time participation of a trusted third party. The protocol requires only two handshakes, significantly reducing the number of protocol rounds compared to related protocols. Finally, the PUF–TRNG architecture has been implemented on the Xilinx XC7A100T development board. Experimental results show that the PUF circuit effectively resists various modeling attacks. Formal verification with ProVerif demonstrates confidentiality, mutual authentication, and robustness against mainstream attacks. The protocol reduces area overhead and computational time by 43.04% and 42.99%, respectively, compared to similar protocols.
Yuanfeng Xie, Weiwei Jiang 0003, Hanqing Luo, Junhong Gan
IEEE Internet Things J.3
2026 A Markov-Chain-Based PUF Using Chain-Block-Obfuscation Mechanism Resisting Machine Learning Attacks With High Uniformity Robustness
abstract
Physical Unclonable Functions (PUFs) are lightweight hardware security primitives suitable for resource-constrained Internet of Things(IoT) devices. The Arbiter PUF (APUF), as a classic strong PUF structure, is well-suited for lightweight device authentication. Unfortunately, due to certain structural characteristics, the classic APUF can be successfully modeled by various machine learning (ML) models with only a small number of CRP (Challenge-Response Pair) samples. To address this security issue, In this paper, we build upon the classic APUF structure by introducing a Chain Block Obfuscation(CBO) mechanism and a Markov obfuscation mechanism at the input and output stages, respectively. This approach demonstrates strong resistance against four machine learning models—logistic regression (LR), support vector machine (SVM), covariance matrix adaptation evolutionary strategies (CMA-ES), and artificial neural networks (ANN)—while optimizing hardware resource usage by at least 53.4% compared to previous attack-resistant structures. Even with up to 2M CRPs for training, the prediction accuracy remains around 50%. Additionally, XOR, a widely adopted output obfuscation mechanism by many researchers, has limitations in output uniformity when the number of intra-chip parallel APUFs is small. Studies have shown that the uniqueness between intra-chip APUFs may not reach the ideal value of 50%. In such cases, especially when there are only two parallel APUFs in the circuit, the uniformity of the XOR output is affected by insufficient uniqueness. To address this issue, this paper proposes a MUX-based Markov output obfuscation mechanism. Experimental data on the Xilinx Artix-7 field-programmable gate array (FPGA) platform demonstrate that this mechanism exhibits strong robustness in uniformity compared to the traditional XOR output obfuscation, particularly when the number of intra-chip parallel APUFs is small.
Junhong Gan, Hanqing Luo, Yuanfeng Xie, Liping Liang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
abstract
This paper introduces SpeedLLM, a neural network accelerator designed on the Xilinx Alevo U280 platform and optimized for the Tinyllama framework to enhance edge computing performance. Key innovations include data stream parallelism, a memory reuse strategy, and Llama2 operator fusion, which collectively reduce latency and energy consumption. SpeedLLM's data pipeline architecture optimizes the read-compute-write cycle, while the memory strategy minimizes FPGA resource demands. The operator fusion boosts computational density and throughput. Results show SpeedLLM outperforms traditional Tinyllama implementations, achieving up to 4.8× faster performance and 1.18× lower energy consumption, offering improvements in edge devices.
Wu Guan, Liping Liang 0001, Hanqing Luo
HPDC5
2025 Design of a high-stability QPUF and QRNG circuit based on CCNOT gate
Yuanfeng Xie, Hanqing Luo, Aoxue Ding
Comput. Secur.2
2025 A Novel High-Throughput FFT Processor With a Block-Level Pipeline for 5G MIMO OFDM Systems
abstract
In fifth-generation (5G) communication systems, multiple input multiple output (MIMO) and orthogonal frequency-division multiplexing (OFDM) are two critical technologies. Fast Fourier transform (FFT), as the core processing steps of OFDM, directly affects the overall system performance. In this brief, we proposed a novel block-level pipelined architecture, which divides the FFT processor into three pipeline blocks: input, radix, and output. Each pipeline block can run in a different FFT simultaneously to achieve higher throughput. Specifically, to reduce the OFDM system-level latency of 5G applications, the FFT processor supports weighted overlap and add (WOLA) on the cyclic prefix and suffix of OFDM symbols. This architecture is implemented using TSMC 12-nm technology, with a processor die area of 0.89 mm2and a power consumption of 568 mW at 1 GHz. The FFT processor can achieve a system-level throughput up to 2.66 GS/s.
Hanqing Luo, Shengnan Lin, Liping Liang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2023 A 60-Mode High-Throughput Parallel-Processing FFT Processor for 5G/4G Applications
abstract
This article presents a 60-mode high-throughput parallel-processing memory-based fast Fourier transform (FFT) processor for fifth-generation (5G)/4G applications. The proposed architecture adopts a multi-FFT parallel-processing scheme to significantly reduce the idle computation cycles of small point FFT due to the deep pipeline. The proposed scheme can narrow the throughput gap between different FFT sizes. In conjunction with the parallel processing scheme, a configurable 16-parallelism butterfly unit is proposed to support a maximum of five radix-3, four radix-4, three radix-5, two radix-8, or one radix-9 operation in one cycle. Furthermore, this article demonstrates conflict-free memory access with a parallel method based on fusion shift combined with a hopping-based method to simplify the data arrangement at different radix stages. According to the orthogonal frequency division multiplexing (OFDM) application characteristics, our processor is extended to support many additional commonly used functions, such as zero padding, cyclic prefix insertion, and 7.5-kHz frequency shift, which can simplify the system call in 5G/4G applications. The FFT processor has been successfully integrated into a small-cell base station baseband system-on-chip (SoC) and taped out at the TSMC 12-nm technology. The implementation result reveals that the die area of the proposed processor is 0.374 mm2 with a power consumption of 235.5 mW at 1 GHz, and the processor supports a balanced throughput up to 3.92 GS/s at all 60-mode, which is better than that of the state-of-the-art (SOTA) designs.
Qinzhi Hong, Hanqing Luo, Xin Qiu 0008, Liping Liang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2012 A wireless force measurement system for Total Knee Arthroplasty
abstract
A wireless force measurement system is presented in this paper. The system, which is used during the operation of Total Knee Arthroplasty(TKA), is designed to assist surgeons to determine the proper position of the knee implants and consequently enhance the success ratio of the treatment. It consists of three parts: a device to measure and transmit force data, a receiver and a terminal to display the force data in real time. The transmitter communicates with the receiver by 2.4GHz Radio Frequency (RF) signal. The system consumes no more than 17mA current with 3V voltage supply typically. So it can work with a button battery cell. Experimental results show that the performance of the system meets the requirements.
Hanqing Luo, Ming Liu 0015, Hong Chen 0002, Chun Zhang 0001, Zhihua Wang 0001
ISCAS1