VLDB 2026 Research / reviewers in the wild / expert
Yijun Cui
dblp:160/0820
· DBLP profile ↗
33ranked-venue papers
6as first author
27since 2021 · last 2026
0000-0002-6262-2329ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 6 first-author · 23 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph-Structure-Aware Hyperdimensional Computing for Hardware Trojan Detection
Zilong Su, Fei Lyu 0002, Yongjun Xia, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2026 | An Efficient Fully-Pipelined Hardware Architecture for Optimized Sparse Polynomial Multiplication in CRYSTALS-Dilithium
Bei Wang 0013, Zeren Zhu, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2026 | Optimized NTT Architecture Based on the Plantard Algorithm for ML-KEM and ML-DSAabstractModular multiplication is a vital operation in the Number Theoretic Transform (NTT), significantly enhancing polynomial multiplication in Post-Quantum Cryptography (PQC). The design efficiency of modular multiplication directly influences the computational performance of polynomial computation units. This work marks the first hardware-oriented improvement of the Plantard algorithm, optimizing the NTT architecture. We modify the Plantard algorithm and propose three innovative enhanced versions tailored for lattice-based cryptography (LBC). By employing pre-processed twiddle factors for result correction and eliminating an additional constant multiplication, we greatly simplify the computation steps. Based on these improvements, we further design a lightweight BRAM-free iterative NTT and a high-speed Multi-path Delay Commutator (MDC) pipelined NTT, both targeting the ML-KEM and ML-DSA parameter sets. Implementation results on the Xilinx Artix-7 platform demonstrate that our Plantard_preω design reduces slice usage by 22% to 53% and delay by 41.1% to 44.3% compared to existing algorithms like Barrett and K2RED. Furthermore, our iterative NTT design achieves the minimal area-time product (ATP) among state-of-the-art implementations, with reductions of 61.3% and 40.8% in ENS, and frequency increases of 72.7% and 145.4% for ML-KEM and ML-DSA, respectively. The pipelined NTT design also reduces delay by 23.7% and ATP by 4.9%, showcasing the compactness and superior performance of our approach. Bei Wang 0013, Ziying Ni, Mengxue Li, Fei Lyu 0002, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 6 |
| 2026 | Instant-CIM: An Instant Neural Radiance Field Computing-In-Memory Architecture for Low-Power and Real-Time AR/VR RenderingabstractNovel View Synthesis is a foundational technique for creating immersive Augmented and Virtual Reality (AR/VR) experiences, aiming to generate photorealistic images of a scene from arbitrary camera viewpoints using only a limited set of source images, with Neural Radiance Fields (NeRF) emerging as the state-of-the-art solution. However, real-time NeRF rendering on low-power devices remains challenging due to its memory-intensive hash encoding and compute-intensive Multilayer Perception (MLP). In this work, we propose Instant-CIM, the fully on-chip Computing-in-Memory (CIM) architecture for efficient NeRF rendering. At the algorithm level, Instant-CIM proposes a spatially-adaptive framework that dynamically selects the number of active hash encoding levels per spatial region based on a composite importance score derived from density and gradient. The approach replaces uniform level allocation with a threshold-based strategy that activates finer encoding levels only in regions with high representation complexity. At the hardware level, Instant-CIM proposes an in-situ hash engine that implements in-memory hash query and interpolation through 3D scene grid decomposition and Z-order based mapping schemes. Meanwhile, Instant-CIM proposes a sparse MLP engine that leverages differential-based input complemented by a precision-adjustable skipping mechanism to fully exploit spatial similarities. Comprehensive evaluation across synthetic datasets demonstrates that Instant-CIM achieves 3.0×~4.9× improvement in rendering speed and 8.6×~33× enhancement in energy efficiency compared to state-of-the-art NeRF architecture. Lixia Han, Hui Chen 0015, Xueming Fu, Ke Chen 0018, Peng Huang 0004, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 8 |
| 2026 | LightHD: A Lightweight and High-Performance Hardware Accelerator of CRYSTALS-DilithiumabstractCRYSTALS-Dilithium serves as the foundation of the NIST-standardised PQC digital signature scheme, and has been declared as the first recommended digital signature algorithm. However, due to the computational complexity and intricate processing flow of CRYSTALS-Dilithium, two limitations are shown in existing methods: its applicability on resource-constrained devices is limited and the performance reported so far remains relatively low. This paper presents a lightweight yet high-performance hardware architecture that optimises the core computational units of CRYSTALS-Dilithium. First, an iterative dual-Keccak SHA-3 module is proposed, where two cores operate with a 26-cycle offset to accelerate processing without compromising frequency. In addition, the rejection sampler is streamlined by two compact registers for intermediate values and counters, improving efficiency when consuming interleaved SHA-3 outputs. Second, for small bit-width polynomials, we eliminate the first NTT stage via lookup tables and data regrouping, reducing NTT cycles by 11.7% with little hardware overhead. Further hardware savings are achieved by maximising IP core utilisation and simplifying input multiplexers. Furthermore, a compact scheduling strategy ensures that all intermediate storage fits within a single polynomial-sized memory block. On Xilinx Artix-7 FPGAs, the design reduces hardware overhead by 14.2% compared with state-of-the-art lightweight implementations. Across three security levels, KeyGen and Verify are 27.3% and 13.5% faster, respectively, than high-performance prior designs. At level 5, the best-case Sign latency is only 120 µs. Ziying Ni, Ayesha Khalid, Zhaoyu Zhang 0001, Yijun Cui, Weiqiang Liu 0001, Máire O'Neill |
IEEE Trans. Computers | 4 |
| 2026 | Approximate Computing-Based Framework for Low-Cost Runtime Hardware Trojan DefenseabstractWith the growing reliance on third-party intellectual property (3PIP) in integrated circuit design, the threat of functional failures induced by hardware Trojans embedded within these components has become increasingly critical. The high stealthiness and sophistication of hardware Trojans often render existing detection techniques insufficient for comprehensive coverage. As a result, runtime detection and recovery techniques have emerged and been proposed as a crucial last line of defense. However, these solutions typically introduce substantial hardware overhead, limiting their practicality in resource-constrained applications such as edge computing. To address this critical challenge, this work leverages approximate computing to develop low-cost runtime recovery strategies against hardware Trojans. Specifically, it begins by analyzing the challenges and potential opportunities introduced into existing security schemes when approximate computing is applied. Based on this analysis, targeted solutions and optimization techniques are proposed. These are then integrated into a unified framework that explores the trade-off between hardware resource usage and computational accuracy while maintaining circuit-level security. Experimental results across several commonly used applications demonstrate that the proposed framework can achieve over 20% of hardware resource savings with only a 5% reduction in accuracy. To the best of our knowledge, this work is the first to systematically integrate approximate computing with runtime hardware Trojan recovery, providing a new cost-effective direction for circuit-level security in resource-constrained systems. Yuqin Dou, Yang Wang 0142, Shiquan Liu, Haroon Waris, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | A High-Accuracy MRAM-Based Computing-in-Memory Macro for Secure Edge AI InferenceabstractComputing-in-memory (CIM) represents a pivotal technology for overcoming the speed and power bottlenecks posed by the “memory wall” and “power wall” existing in von Neumann architectures. With the fast development of Internet of Things, data security has become one of the most attractive research topics for CIM in edge applications, as well as computing accuracy and energy efficiency. This paper proposes a highly accurate and secure MRAM-based CIM macro that is designed to reduce the multiply accumulate (MAC) computation errors and protect weight bits in untrusted environments. The architecture employs a series-connected structure to enhance computational linearity and introduces a dynamic reference column to increase reliability. Meanwhile, a lightweight encryption mechanism based on physical unclonable function (PUF) is implemented to protect weight bits. The results demonstrate that, after the obfuscation, the prediction accuracy of machine learning-based attacks on the PUF is reduced to approximately 50%. The CIM macro achieves impressive inference accuracy of 93.73% on the CIFAR-10 dataset with energy efficiency of 38.3 TOPS/W. Furthermore, security verification performed on the CNN model indicates that weight encryption degrades inference accuracy to 10%, thereby providing robust protection against potential attacks. You Wang 0002, Jiaao Dai, Shuo Fan, Yijun Cui, Yu Gong 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Exploring Teaching Methods for Courses on Radiation Hardening Technology in ICsabstractWith the rapid advancement of space exploration technology, the use of intelligent equipment and systems is increasing at an accelerated pace. As the core component of intelligent systems, integrated circuits (ICs) have become a key area of research in space applications. However, the complex space environment significantly degrades the reliability of ICs due to radiation effects. As a result, radiation hardening technology is critical for ICs used in space applications. Unlike general consumer electronics, students majoring in ICs are often unfamiliar with radiation hardening technologies, which is a disadvantage for those who may work in industries such as aerospace, nuclear, or medical electronics after graduation. This paper explores teaching methods for a course on radiation hardening technology in ICs. Through interdisciplinary collaboration and joint university-enterprise teaching, as well as classroom interaction and project-based learning, students will gain an in-depth understanding of radiation sources, radiation effects, hardening techniques, and irradiation testing. You Wang 0002, Erya Deng, Yu Gong 0002, Zhongkun Shen, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 6 |
| 2025 | An Efficient Hardware Implementation of Improved Plantard Mod-Multiplication for Lattice-Based CryptographyabstractThe modular multiplication (mod-multiplication) algorithm is an essential operation in lattice-based cryptography (LBC) that utilizes the Number Theoretical Transform (NTT) for polynomial multiplication. An efficient mod-multiplication algorithm determines the computational efficiency and performance of the entire polynomial multiplier/NTT computation unit. In this manuscript, we propose an improved Plantard modmultiplication algorithm for the NTT in Kyber and Dilithium, which not only reduces one multiplication but also eliminates the post-processing operations compared with the original Plantard algorithm. Additionally, we design an optimized hardware implementation for the improved Plantard mod-multiplication algorithm. Based on the Xilinx Artix-7 platform, when compared with state-of-the-art designs, our improved Plantard algorithm reduces the number of slices by 22.2%∼53.3% for Kyber and 18.4%∼35.4% for Dilithium, while boosting hardware efficiency by 43.9%∼55.9% for Kyber and 33.7%∼54.4% for Dilithium. Overall, our improved Plantard algorithm shows significant advantages in resource consumption and computational speed. Mengxue Li, Bei Wang 0013, Fei Lyv, Weiqiang Liu 0001, Yijun Cui |
ISCAS | 6 |
| 2025 | Ultra-compact and Side-channel Resistant Design of FIFO-based NTT Core for PQCsabstractCryptographic algorithms like CRYSTALS-Kyber and Dilithium might be insecure with their naive implementation facing side-channel attacks (SCA). This work presents a compact implementation of Number Theoretic Transform (NTT) with shuffling countermeasure against power analysis attacks (PA). At first, a compact FIFO-only Shuffler module is presented to perform group-wise first-index randomization (FIR). A modified butterfly (BF) unit using optimized modulus reduction is then promoted to restore the misaligned data flow, which is critical for forming efficient shuffle pattern. The shuffler module and BF unit are then used in a pipelined BRAM-free NTT baseline. Through efficient shuffling, the proposed design maintains compactness akin to its baseline while enhancing robust SCA resistance with a permutation space of up to 2494. Compared to its prior state-of-the-art designs, the proposed NTT core presents an improvement of 53.2% in area-time trade-off while offering ×1.7 times more bits of randomness to improve hardware security. Jiatong Tian, Yijun Cui, Ziying Ni, Bei Wang 0013, Fei Lyv, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2025 | A Lightweight and Efficient BRAM-free NTT Unit for Crystals-DilithiumabstractDuring the standardization of post-quantum cryptography by the National Institute of Standards and Technology (NIST), the lattice-based Crystals-Dilithium algorithm was selected as the standardized digital signature scheme. This work designs a lightweight and efficient BRAM-free Number Theoretic Transform (NTT) unit, which is a major bottleneck for Crystals-Dilithium. Firstly, we propose an improved parallel modular multiplication based on the K-RED algorithm, effectively reducing resource consumption and shortening the critical path. Furthermore, a BRAM-free iterative NTT architecture is designed, utilizing three first-in-first-out (FIFO) buffers to store intermediate data. Evaluated on the Xilinx Artix-7 and Zynq UltraScale+ platforms, our proposed NTT architecture presents the best hardware efficiency with less resource consumption. Experimental results show that our design is 36.7%-80.9% reduced in terms of resource consumption and is 31.5%-89.9% better in terms of hardware efficiency compared with state-of- the-art works. Junjie Zhong, Bei Wang 0013, Zeren Zhu, Weiqiang Liu 0001, Yijun Cui |
ISCAS | 5 |
| 2025 | High-Performance Hardware Implementation of Crystals-Dilithium Based on Improved MDC-NTTabstractThe growing threat of quantum computing to traditional cryptographic systems has necessitated the development of robust post-quantum algorithms. Crystal-Dilithium, recently standardized by NIST after a three-round competition, is a leading lattice-based digital signature algorithm designed to meet this need. However, conventional hardware implementations of Dilithium often suffer from inefficiencies and performance bottlenecks. To address these weaknesses, this work presents an optimized hardware design for Dilithium across all security levels. The proposed design features a parallel modular multiplication unit, and an enhanced scaling method to reduce bit width and minimize calibration. Additionally, an improved radix-2 Multipath Delay Commutator Number Theoretic Transform (MDC-NTT) and pipelined parallelization using FIFO and BRAM-based buffers are integrated to maximize operating frequency. Evaluated on the Xilinx Artix-7 platform, our implementation achieves a peak frequency of 191 MHz, delivering speedups of 26.3%, 32.5% and 29.6% for key generation, signature generation and signature verification respectively, compared with state-of-the-art works at the highest security level, along with superior hardware efficiency. Yijun Cui, Junjie Zhong, Bei Wang 0013, Tianyu Xu 0002, Chenghua Wang, Weiqiang Liu 0001 |
IEEE Trans. Computers | 1 |
| 2025 | A Highly Reliable Dual-Mode RRAM PUF With Key Concealment SchemeabstractPhysical unclonable function (PUF) has been widely used in the Internet of Things (IoT) as a promising hardware security primitive. In recent years, PUFs based on resistive random access memory (RRAM) have demonstrated excellent reliability and integration density. Most previous designs store PUF keys directly in RRAMs, increasing vulnerability to attacks. This article proposes a dual-mode RRAM PUF, named differential mode and flexible mode, utilizing the difference in switching capability between RRAMs during parallel SET operations as the entropy source. The proposed PUF can reliably reproduce keys between cycles, so a key concealment scheme is used to protect PUF keys from being continuously exposed, improving the security of the RRAM PUF. The proposed RRAM PUF exhibits high reliability over ±10% VDD and a wide temperature range from −25°C to 125°C through post-processing operations. The flexible mode can generate a significant number of keys for high-security applications. Since the PUF keys can be concealed, the proposed PUF is compatible with in-memory computing. It can be implemented using the same RRAM array as experimentally validated using a MAGIC operation, thus reducing the hardware overhead. Jiang Li 0012, Yijun Cui, Chongyan Gu, Chenghua Wang, Weiqiang Liu 0001, Shahar Kvatinsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Instruction-Based High-Performance Hardware Controller of CRYSTALS-Kyber With Balanced Resource UtilizationabstractPost-quantum cryptography (PQC) aims to ensure information security in the era following the emergence of quantum computers. Lattice-based cryptography (LBC) algorithms have shown significant promise in the standardization process of post-quantum cryptography. This paper proposes an instruction-based high-performance hardware controller of CRYSTALS-Kyber. By designing a highly flexible instruction-based architecture, the control unit evenly distributes instructions and enables independent control of internal modules, significantly enhancing the scalability and adaptability of the hardware. Additionally, the integration of a reconfigurable polynomial operation array (RPOA) unit and optimization of data storage formats further improve computational efficiency and resource utilization. Implementation results on Artix-7 FPGA show that the architecture operates at a frequency exceeding 300 MHz, achieving a performance improvement of 41.3% to 170% compared to the latest designs, while significantly reducing resource overhead. The resource costs for the three security levels are 8112 LUTs, 6077 FFs, and 2523 SLICEs, respectively, with overall computation times of$34.7~\mu s$,$53.4~\mu s$, and$78.5~\mu s$. The proposed design demonstrates outstanding performance, resource efficiency, and energy consumption, providing an efficient and cost-effective hardware solution for the practical deployment of post-quantum cryptography. Yijun Cui, Ziying Ni, Zhuoyao Zhang, Chenghua Wang, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | Optimal Linear Codes From Duals of Punctured Concatenated CodesabstractA code is called a punctured concatenated code if it can be obtained by puncturing a concatenated code at suitable coordinates. Based on this new concept, we construct several classes of optimal or almost optimal linear codes. There are two major contributions in this paper. Let the inner code be an [n,m]qlinear code derived from the defining setD= {d1,d2, . . . ,dn}. On the one hand, by employing a maximum distance separable (MDS) code with dimension 2 over Fqmas the outer code, we propose two classes of linear codes with few weights. The duals of these codes are shown to be dimension-optimal with respect to the sphere-packing bound. On the other hand, letq= 2, by choosing an MDS code with dimension 3 over F2mas the outer code, we construct another class of linear codes. The parameters and weight distributions of these codes are completely determined. Furthermore, their dual codes are almost distance-optimal with respect to the sphere-packing bound. Gaojun Luo, Yijun Cui, Xiwang Cao, San Ling |
IEEE Trans. Inf. Theory | 3 |
| 2025 | A 0.09-pJ/Bit Logic-Compatible Multiple-Time Programmable (MTP) Memory-Based PUF Design for IoT ApplicationsabstractThe Internet of Things (IoT) allows devices to interact for real-time data transfer and remote control. However, IoT hardware devices have been shown security vulnerabilities. Edge device authentications, as a crucial process for IoT systems, generate and use unique IDs for secure data transmissions. Conventional authentication techniques, computational and heavyweight, are challenging and infeasible in IoT due to limited resources in IoTs. Physical unclonable functions (PUFs), a lightweight hardware-based security primitive, were proposed for resource-constrained applications. We propose a new PUF design for resource-constrained IoT devices based on low-cost logic-compatible multiple-time programmable (MTP) memory cells. The structure includes an array of MTP differential memory cells and a PUF extraction circuit. The extraction method uses the random distribution of BL current after programming each memory cell in logic-compatible MTP memory as the entropy source of PUF. Responses are obtained by comparing the current values of two memory cells under a certain address by challenge, forming challenge–response pairs (CRPs). This scheme does not increase hardware consumption and circuit differences on edge devices and is intrinsic PUF. Finally, 200 PUF chips were fabricated by CSMC based on the 0.153-$\mu $m MCU single-gate CMOS process. The performance of the logic-compatible MTP memory cell and its PUF was evaluated. A logic-compatible MTP cell has good programming erase efficiency and good durability and retention. The uniqueness of the proposed PUF is 50.29%, the uniformity is 51.82%, and the reliability is 93.61%. Shuming Guo, Yinyin Lin, Yao Li 0018, Chongyan Gu, Weiqiang Liu 0001, Yijun Cui |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | A Concealable RRAM Physical Unclonable Function Compatible with In-Memory ComputingabstractResistive random access memory (RRAM) has been widely used in physical unclonable function (PUF) design due to its low power consumption, fast read/write speed, and significant intrinsic randomness. However, existing RRAM PUFs cannot overcome the cycle-to-cycle (C2C) variations of RRAM, leading to poor reproducibility of PUF keys across cycles. Most prior designs directly store PUF keys in RRAMs, increasing vulnerability to attacks. In this paper, we propose a concealable RRAM PUF based on an RRAM crossbar array, utilizing the differential resistive switching characteristics of two RRAMs to generate keys. By enabling the reproducibility of PUF keys across cycles, a concealment scheme is proposed to prevent the exposure of PUF keys, thus enhancing the security of the RRAM PUF. Through post-processing operations, the proposed PUF exhibits high reliability over ±10% VDD and a wide temperature range from 248K to 373K. Furthermore, this RRAM PUF is compatible with in-memory computing (IMC), and they can be implemented using the same RRAM crossbar array. Jiang Li 0012, Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Shahar Kvatinsky |
DATE | 2 |
| 2024 | A Novel Methodology for Processor based PUF in Approximate ComputingabstractApproximate computing has great potential in the design of high-performance and energy-efficient systems. The inherent stochastic error behavior of approximate computing introduces both new security threats and opportunities to enhance security. This work proposes a novel methodology that exploits stochastic timing errors of a pipelined datapath to design a processor based physically unclonable function (PUF) for approximate computing. This methodology uses divergent delay path selection based on intermediary error behaviour to improve the PUF uniqueness vs. an unmodified datapath, even when only moderate voltage scaling is applied. To verify the effectiveness of this method, a pipelined fast fourier transform (FFT) butterfly architecture is implemented at 45nm technology node, and a voltage over scaling technique is applied to extract PUF bits. The proposed methodology achieves a maximum uniqueness of 48.5% whereas conventional design uniqueness is limited to 43%. Overall, the proposed design shows a maximum of ~7% higher uniqueness and ~10% higher reliability (for iso uniqueness) compared to the conventional pipelined design. Aditya Japa, Jack Miskelly, Yijun Cui, Máire O'Neill, Chongyan Gu |
ISCAS | 3 |
| 2024 | Lattice-based Multi-Stage Secret Sharing 3D Secure Encryption SchemeabstractWith the widespread deployment of three-dimensional (3D) models in industry and daily life, protecting the security of this data becomes crucial. Additionally, three-dimensional (3D) models may be distributed to users with varying security levels, necessitating distinct visualizations for each user. Recent research proposes 3D model encryption method that facilitates distinct visualizations post-decryption through hierarchical decryption. However, this method permits the decryption of 3D models at varying visual security levels based on user privileges. It has potential security vulnerabilities concerning key management and simultaneously limits its capacity to address diverse user requirements. To address this, a multi-stage secret sharing mechanism is integrated into the existing hierarchical encryption framework to bolster the security of hierarchical keys. When combined with lattice-based cryptography techniques, it ensures that only users with adequate shares can decrypt the corresponding 3D model hierarchy, achieving distinct visual effects while maintaining secret key security under diverse user needs. Experimental results demonstrate that the scheme effectively enhances security while maintaining data integrity and availability. Yinghao Wu, Bei Wang 0013, Yijun Cui |
TrustCom | 6 |
| 2024 | An Efficient Ring Oscillator PUF Using Programmable Delay Units on FPGAabstractThe ring oscillator (RO) PUF can be implemented on different FPGA platforms with high uniqueness and reliability. To decrease the hardware cost of conventional RO PUFs, a new design using the programmable delay units is proposed, namely, PRO PUF. The programmable interconnect points (PIPs) of programmable delay units are used to enhance the configurability. The PUF cell of the proposed design has the ability to be efficiently programmed to an RO PUF at any stage by adjusting the propagation paths of the delay units. A significant number of responses can be generated by the proposed PRO PUF while consuming fewer hardware resources. To verify the performance, the proposed design has been implemented on Xilinx FPGAs and also simulated using a standard 40nm technology. The experimental results have shown that the proposed design achieves high uniqueness, reliability, and hardware efficiency. Moreover, the PRO PUF has been evaluated using a machine learning attack, the CMA-ES attack. The results have shown that the proposed structure is more resistant to common modeling attacks when compared to conventional RO-related PUF designs. Yijun Cui, Jiang Li 0012, Yunpeng Chen, Chenghua Wang, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2023 | Novel Intrinsic Physical Unclonable Function Design for Post-quantum CryptographyabstractThe hardware implementations of post-quantum cryptography (PQC) algorithms are vulnerable to fault injection attacks. As a hardware security primitive, the intrinsic physical unclonable function (PUF) is a possible countermeasure for these attacks with low resource overheads. In this work, a novel intrinsic PUF, frequency adjustable software PUF (FAS-PUF), is proposed to provide a device identification for PQC chips. The FAS-PUF is based on an inherent timing logic in the ring-learning with error (R-LWE) decryption circuit of PQC chips. The FAS-PUF uses a$256^{\ast}13^{\ast} 3$-bit input ciphertext of the decryption circuit as a challenge, and uses a 256-bit decryption output as a response with an adjustable overclocking. Since the entropy of the FAS-PUF utilises the manifested timing errors caused by the overclocking, the FAS-PUF does not need to modify the existing hardware circuits, i.e. preserves the original circuit functions, which significantly reduces hardware resource consumption and power overhead. Meanwhile, to mitigate the affection of circuits' metastablities to PUF's stability under overclocking, a dynamic clock frequency selection method is used to determine the optimal frequency point for generating PUF responses. The proposed FAS-PUF is also a Strong PUF design with a significant number of Challenge/Response Pairs (CRPs) provided. The proposed design is implemented on Xilinx Basys3 FPGAs. The experimental results show that the FAS-PUF has a good uniqueness, uniformity and stability compared with other intrinsic PUFs. Yijun Cui, Chongyan Gu, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2023 | How Practical Phase-Shift Errors Affect Beamforming of Reconfigurable Intelligent Surface?abstractReconfigurable intelligent surface (RIS) is able to manipulate the wireless environment smartly and has been exploited for assisting the wireless communications, especially at high frequency band. However, it suffers from hardware impairments (HWIs) in practical design, manufacturing and deployment, which inevitably degrades its performance and thus limits its full potential. To address this practical issue, we first propose a new RIS reflection model involving phase-shift errors, which is verified by the measurement results from field trials. With this beamforming model, various phase-shift errors caused by different HWIs can be analyzed. The phase-shift errors are classified into three categories: 1) globally independent and identically distributed errors; 2) grouped independent and identically distributed errors; and 3) grouped fixed errors. The impact of typical HWIs, including frequency mismatch, PIN diode failures and panel deformation, on RIS beamforming ability are studied with the theoretical model and are verified with numerical and field test data. The impact of frequency mismatch are discussed separately for narrow-band and wide-band beamforming. Finally, useful insights and guidelines on the RIS design and its deployment are highlighted for practical wireless sytsems. Jun Yang 0058, Yijian Chen, Yijun Cui, Qingqing Wu 0001, Jianwu Dou |
IEEE Trans. Commun. | 3 |
| 2022 | Horizontal Correlation Analysis without Precise Location on Schoolbook Polynomial Multiplication of Lattice-based CryptosystemabstractMost cryptographic systems are secure in theory; however, the implementation of cryptographic system on embedded devices can be attacked by analyzing the power consumption of specific operation to reveal the key. The classic vertical correlation power analysis (CPA) attack requires a large number of power traces for analysis. Using transient secret-key scheme significantly weakens such an attack as insufficient data could be obtained. On the other hand, the horizontal CPA requires at least a single power trace and can make full use of multiple intermediate values to analyze the correlation of power consumption. In this work, we devised a horizontal CPA attack on schoolbook polynomial multiplication of hardware-implemented lattice-based cryptosystem without precise location. The accuracy of correctly recovering any one sub secret-key using only a single trace is 99.90%, and the accuracy of correctly recovering the secret-key is 76.41%. The powerful attack capability of horizontal CPA exposes the vulnerability of unprotected schoolbook polynomial multiplication against the attack of side-channel analysis (SCA). Chuanchao Lu, Yijun Cui, Dur-e-Shahwar Kundi, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2022 | A Lightweight and Efficient Schoolbook Polynomial Multiplier for SaberabstractSaber is a lattice-based post-quantum cryptography (PQC) algorithm, which is still a candidate in the 3rdRound of National Institute of Standards and Technology (NIST) PQC standardization process. Saber provides a great advantage of being lightest among all the candidates, so a suitable choice for resource-constraint platforms. Polynomial multiplication occupies most of the resources in hardware implementation of Saber, which needs to be optimized for the efficient hardware implementation. In this work, a lightweight and efficient schoolbook polynomial multiplier is proposed. The architecture includes an efficient multiplication strategy that compute four coefficient-wise multiplication per cycle along with the multiplication operand loading technique being designed for the compact multiplier. The proposed multiplier on Artix-7 FPGA, achieves a frequency of 130 MHz and fits into 201 slices. Compared with the state-of-the-art lightweight schoolbook implementations for Saber, our design has a 30% improved frequency and saves 15.8% of the clock counts at the cost of only 3.7% more LUTs. Yuantuo Zhang, Yijun Cui, Ziying Ni, Dur-e-Shahwar Kundi, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2022 | A Generic Dynamic Responding Mechanism and Secure Authentication Protocol for Strong PUFsabstractAs a lightweight hardware security primitive, physical unclonable functions (PUFs) can provide reliable identity authentication for devices of Internet of Things (IoTs) with limited resources. However, the delay-based PUF structures in authentication protocols have static responding behaviors, which make them vulnerable to modeling attacks. To address this issue, many complex PUF designs have been designed to increase the nonlinearity of their models. However, most of them can still be broken by modeling-based machine learning (ML) attacks. In this article, a dynamic responding mechanism for PUF designs to generate dynamic responses is proposed. Different from the concept of logically reconfigurable PUFs, the proposed mechanism does not rely on external inputs to provide reconfiguration signals. And different from the conventional PUF authentication protocols that use large-size linear feedback shift register (LFSR) to extend the master challenge, the proposed scheme uses internally generated dynamic signals to obfuscate the master challenge to generate multiple subchallenges. These subchallenges are then input to the underlying strong PUF to generate multibit dynamic responses. It can prevent an attacker from obtaining valid challenge-response pairs (CRPs) for the underlying PUF. A security authentication protocol is also proposed, the special authentication bit-string design can resist both conventional ML attacks and the latest covariance matrix adaptation evolution strategies (CMA-ES) variant. Yale Wang, Chenghua Wang, Chongyan Gu, Yijun Cui, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | A Dynamic Highly Reliable SRAM-Based PUF Retaining Memory FunctionabstractIn this paper, a highly reliable SRAM based Physical Unclonable Function (PUF), which retains the memory function is proposed. The mismatch of NMOS is extracted during discharge process and amplified by the cross-coupled inverter to generate a response. At the beginning of the discharge process, the NMOSs are biased at sub-threshold region, which can improve the reliability and stability. The proposed PUF is designed in a 40nm CMOS process and each bit cell only consumes 4.98 μm2(3112F2). Post simulation shows that the bit error rate (BER) deterioration is 0.96% per 0.1V, 0.36% per 10° C with temperature variations from -40° C to 80° C and supply voltage variations from 0.9V to 1.3V. It achieves 1.8% native instability through the simulation. Meanwhile, the proposed PUF can retain memory function after a response is generated. Chenghua Wang, Chenggang Yan 0002, Yijun Cui, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2021 | PUF-Based Mutual-Authenticated Key Distribution for Dynamic Sensor NetworksabstractBecause of the movements of sensor nodes and unknown mobility pattern, how to ensure two communicating (static or mobile) nodes authenticate and share a pairwise key is important. In this paper, we propose a mutual-authenticated key distribution scheme based on physical unclonable functions (PUFs) for dynamic sensor networks. Compared with traditional key predistribution schemes, the proposal reduces the storage overhead and the key exposure risks and thereby improves the resilience against node capture attacks. Mutual authentication is provided by the PUF challenge-response mechanism. However, the PUF response is not transmitted in plain forms so as to resist the modelling attacks, which is vulnerable in some existing PUF-based schemes. We demonstrate the proposed scheme to improve the secure connectivity and other performances by analysis and experiments. Yijun Cui, Lein Harn, Shuo Qiu |
Secur. Commun. Networks | 2 |
| 2020 | Programmable Ring Oscillator PUF Based on Switch MatrixabstractConfigurable ring oscillator (CRO) physical unclonable functions (PUFs) which can improve the uniqueness and reliability of conventional RO PUFs have been widely studied. Especially, the multiplier, XOR gate and tristate inverter based CRO PUFs can improve the uniqueness and reliability. However the efficiency is remain at the same level when compared with the conventional RO PUFs. In this paper, a programmable RO PUF (PRO PUF), which can be programmed to change the structure of a typical RO PUF, is proposed. The proposed PRO PUF design is implemented based on the switch matrix of an FPGA and can be programmed as a chained RO PUF or a random looped RO PUF. The proposed PRO PUF is implemented on Xilinx Spartan 6 FPGAs. Experimental results demonstrate that the proposed PRO PUF design has good uniqueness and reliability metrics as well as a high hardware efficiency. Yijun Cui, Yunpeng Chen, Chenghua Wang, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 1 |
| 2019 | Theoretical Analysis of Delay-Based PUFs and Design Strategies for ImprovementabstractDelay-based physical unclonable function (PUF) designs use the random delay differences in circuit transmission to extract response. In the existing PUF designs, there are few studies on investigating the link between process variation and PUF performance. The experimental data can reflect the performance of the new design to a certain extent, but lack of theoretical analysis to provide thorough information. In this paper, a theoretical model for delay-based PUF designs is proposed. An analysis of the delay-based PUF improvements by existing design strategies is also investigated. Moreover, a guidance to develop and improve future delay-based PUF designs using the proposed theoretical model is also given in this paper. Yale Wang, Chenghua Wang, Chongyan Gu, Yijun Cui, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2019 | Multi-Incentive Delay-Based (MID) PUFabstractThis paper proposes a new PUF, namely Multi-incentive Delay-based PUF (MID PUF), which utilizes the fast carry logic (FCL) of Field Programmable Gate Arrays (FPGAs). The proposed MID PUF is completely and efficiently implemented in XOR gates of FCLs. Compared to other single signal excited PUF designs, e.g. Arbiter PUF, multiple excitations are applied on the same delay line to produce multiple outputs. To the authors' best knowledge, this is the first strong PUF based on only FCLs. The proposed MID PUF is implemented on Xilinx Spartan-6 XC6SLX9 FPGAs and a reliability experiment is carried out under the operating temperature in a range of 0°C~70° C. The experimental results show that the proposed MID PUF has a high uniqueness and reliability performance, as well as low hardware consumption. Due to its advantages in both hardware efficiency and PUF metrics, the proposed MID PUF is promising for low-cost security applications on FPGAs. Zhengran Zhang, Chongyan Gu, Yijun Cui, Chuan Zhang 0001, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2016 | Live demonstration: An automatic evaluation platform for physical unclonable function testabstractPUF is a security primitive that exploits the fact that no two ICs are exactly the same. To verify a new PUF design, several metrics including uniqueness, reliability, and randomness must be evaluated, which requires various resources and a long set-up time. In this live demonstration, we have developed an automatically evaluation platform for the PUF design. To the authors' best knowledge, this is the first automatic evaluation platform for the PUF test. The evaluation platform can be used for both FPGA and ASCI PUF testing. Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 1 |
| 2016 | Low-cost configurable ring oscillator PUF with improved uniquenessabstractThe physical unclonable function (PUF) produces die-unique responses and is regarded as an emerging security primitive that can be used for authentication of devices. The complexity of a conventional PUF design based on a ring oscillator (RO) is rather high, so limiting its use in many applications. The configurable ring oscillator (CRO) PUF has been advocated as a possible solution to this issue. In this paper, a low hardware complexity CRO PUF design with an enhanced capability to generate a large number of bit responses is proposed; only an inverter and a multiplexer are used in each delay unit. The responses are generated by considering the variation due to fabrication of the logic gates and wires in the CROs. A novel comparison strategy is proposed for the generation of the responses. The proposed PUF design is implemented on Xilinx Spartan-6 FPGAs. These results show that the proposed CRO PUF design has good uniqueness; moreover, it is also robust in its operation for the temperature range of -25°C~85°C. Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill, Fabrizio Lombardi |
ISCAS | 1 |
| 2015 | RO PUF design in FPGAs with new comparison strategiesabstractA Physical Unclonable Function (PUF) can be used to provide authentication of devices by producing die-unique responses. In PUFs based on ring oscillators (ROs), the responses are derived from the oscillation frequencies of the ROs. However, RO PUFs can be vulnerable to attack due to the frequency distribution characteristics of the RO arrays. In this paper, in order to improve the design of RO PUFs for FPGA devices, the frequencies of RO arrays implemented on a large number of FPGA chips are statistically analyzed. Three RO frequency distribution (ROFD) characteristics are observed and discussed. Based on these ROFD characteristics, two RO comparison strategies are proposed that can be used to improve the design of RO PUFs. It is found that the symmetrical RO comparison strategy has the highest entropy density. Weiqiang Liu 0001, Chenghua Wang, Yijun Cui, Máire O'Neill |
ISCAS | 4 |