EDBT 2026 Demo / reviewers in the wild / expert
Chenghua Wang
dblp:76/1643
· DBLP profile ↗
45ranked-venue papers
0as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 23 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph-Structure-Aware Hyperdimensional Computing for Hardware Trojan Detection
Zilong Su, Fei Lyu 0002, Yongjun Xia, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2026 | An 8T-SRAM Near-Memory Architecture for Multiplierless Approximate DCT
Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Lixia Han, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 6 |
| 2026 | An Efficient Fully-Pipelined Hardware Architecture for Optimized Sparse Polynomial Multiplication in CRYSTALS-Dilithium
Bei Wang 0013, Zeren Zhu, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2026 | VLMCache: Efficient On-Device Vision-Language Model InferenceabstractVision Language Models (VLMs) are foundational for low-latency, privacy-preserving on-device AI in real-time applications like UI agents and VQA. The VLM prefilling phase, which processes the entire visual-textual input, faces the critical challenge of a long Time-to-First-Token (TTFT). One promising approach to reduce TTFT is to exploit the temporal locality by reusing block-level computations across consecutive frames. Unfortunately, current Transformer-based VLMs break the spatial invariance of CNNs and invalidate the strict-prefix KV-cache mechanism of decoder-only LLMs; in practice, even a single-pixel mismatch can prevent reuse. Yinyuan Zhang, Daliang Xu, Chenghua Wang, Ying Zhang 0012, Mengwei Xu 0001, Gang Huang 0001 |
MobiSys | 4 |
| 2025 | Exploring Teaching Methods for Courses on Radiation Hardening Technology in ICsabstractWith the rapid advancement of space exploration technology, the use of intelligent equipment and systems is increasing at an accelerated pace. As the core component of intelligent systems, integrated circuits (ICs) have become a key area of research in space applications. However, the complex space environment significantly degrades the reliability of ICs due to radiation effects. As a result, radiation hardening technology is critical for ICs used in space applications. Unlike general consumer electronics, students majoring in ICs are often unfamiliar with radiation hardening technologies, which is a disadvantage for those who may work in industries such as aerospace, nuclear, or medical electronics after graduation. This paper explores teaching methods for a course on radiation hardening technology in ICs. Through interdisciplinary collaboration and joint university-enterprise teaching, as well as classroom interaction and project-based learning, students will gain an in-depth understanding of radiation sources, radiation effects, hardening techniques, and irradiation testing. You Wang 0002, Erya Deng, Yu Gong 0002, Zhongkun Shen, Chenghua Wang, Yijun Cui, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2025 | Ultra-compact and Side-channel Resistant Design of FIFO-based NTT Core for PQCsabstractCryptographic algorithms like CRYSTALS-Kyber and Dilithium might be insecure with their naive implementation facing side-channel attacks (SCA). This work presents a compact implementation of Number Theoretic Transform (NTT) with shuffling countermeasure against power analysis attacks (PA). At first, a compact FIFO-only Shuffler module is presented to perform group-wise first-index randomization (FIR). A modified butterfly (BF) unit using optimized modulus reduction is then promoted to restore the misaligned data flow, which is critical for forming efficient shuffle pattern. The shuffler module and BF unit are then used in a pipelined BRAM-free NTT baseline. Through efficient shuffling, the proposed design maintains compactness akin to its baseline while enhancing robust SCA resistance with a permutation space of up to 2494. Compared to its prior state-of-the-art designs, the proposed NTT core presents an improvement of 53.2% in area-time trade-off while offering ×1.7 times more bits of randomness to improve hardware security. Jiatong Tian, Yijun Cui, Ziying Ni, Bei Wang 0013, Fei Lyv, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 6 |
| 2025 | High-Performance Hardware Implementation of Crystals-Dilithium Based on Improved MDC-NTTabstractThe growing threat of quantum computing to traditional cryptographic systems has necessitated the development of robust post-quantum algorithms. Crystal-Dilithium, recently standardized by NIST after a three-round competition, is a leading lattice-based digital signature algorithm designed to meet this need. However, conventional hardware implementations of Dilithium often suffer from inefficiencies and performance bottlenecks. To address these weaknesses, this work presents an optimized hardware design for Dilithium across all security levels. The proposed design features a parallel modular multiplication unit, and an enhanced scaling method to reduce bit width and minimize calibration. Additionally, an improved radix-2 Multipath Delay Commutator Number Theoretic Transform (MDC-NTT) and pipelined parallelization using FIFO and BRAM-based buffers are integrated to maximize operating frequency. Evaluated on the Xilinx Artix-7 platform, our implementation achieves a peak frequency of 191 MHz, delivering speedups of 26.3%, 32.5% and 29.6% for key generation, signature generation and signature verification respectively, compared with state-of-the-art works at the highest security level, along with superior hardware efficiency. Yijun Cui, Junjie Zhong, Bei Wang 0013, Tianyu Xu 0002, Chenghua Wang, Weiqiang Liu 0001 |
IEEE Trans. Computers | 5 |
| 2025 | A Highly Reliable Dual-Mode RRAM PUF With Key Concealment SchemeabstractPhysical unclonable function (PUF) has been widely used in the Internet of Things (IoT) as a promising hardware security primitive. In recent years, PUFs based on resistive random access memory (RRAM) have demonstrated excellent reliability and integration density. Most previous designs store PUF keys directly in RRAMs, increasing vulnerability to attacks. This article proposes a dual-mode RRAM PUF, named differential mode and flexible mode, utilizing the difference in switching capability between RRAMs during parallel SET operations as the entropy source. The proposed PUF can reliably reproduce keys between cycles, so a key concealment scheme is used to protect PUF keys from being continuously exposed, improving the security of the RRAM PUF. The proposed RRAM PUF exhibits high reliability over ±10% VDD and a wide temperature range from −25°C to 125°C through post-processing operations. The flexible mode can generate a significant number of keys for high-security applications. Since the PUF keys can be concealed, the proposed PUF is compatible with in-memory computing. It can be implemented using the same RRAM array as experimentally validated using a MAGIC operation, thus reducing the hardware overhead. Jiang Li 0012, Yijun Cui, Chongyan Gu, Chenghua Wang, Weiqiang Liu 0001, Shahar Kvatinsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Instruction-Based High-Performance Hardware Controller of CRYSTALS-Kyber With Balanced Resource UtilizationabstractPost-quantum cryptography (PQC) aims to ensure information security in the era following the emergence of quantum computers. Lattice-based cryptography (LBC) algorithms have shown significant promise in the standardization process of post-quantum cryptography. This paper proposes an instruction-based high-performance hardware controller of CRYSTALS-Kyber. By designing a highly flexible instruction-based architecture, the control unit evenly distributes instructions and enables independent control of internal modules, significantly enhancing the scalability and adaptability of the hardware. Additionally, the integration of a reconfigurable polynomial operation array (RPOA) unit and optimization of data storage formats further improve computational efficiency and resource utilization. Implementation results on Artix-7 FPGA show that the architecture operates at a frequency exceeding 300 MHz, achieving a performance improvement of 41.3% to 170% compared to the latest designs, while significantly reducing resource overhead. The resource costs for the three security levels are 8112 LUTs, 6077 FFs, and 2523 SLICEs, respectively, with overall computation times of$34.7~\mu s$,$53.4~\mu s$, and$78.5~\mu s$. The proposed design demonstrates outstanding performance, resource efficiency, and energy consumption, providing an efficient and cost-effective hardware solution for the practical deployment of post-quantum cryptography. Yijun Cui, Ziying Ni, Zhuoyao Zhang, Chenghua Wang, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | A Concealable RRAM Physical Unclonable Function Compatible with In-Memory ComputingabstractResistive random access memory (RRAM) has been widely used in physical unclonable function (PUF) design due to its low power consumption, fast read/write speed, and significant intrinsic randomness. However, existing RRAM PUFs cannot overcome the cycle-to-cycle (C2C) variations of RRAM, leading to poor reproducibility of PUF keys across cycles. Most prior designs directly store PUF keys in RRAMs, increasing vulnerability to attacks. In this paper, we propose a concealable RRAM PUF based on an RRAM crossbar array, utilizing the differential resistive switching characteristics of two RRAMs to generate keys. By enabling the reproducibility of PUF keys across cycles, a concealment scheme is proposed to prevent the exposure of PUF keys, thus enhancing the security of the RRAM PUF. Through post-processing operations, the proposed PUF exhibits high reliability over ±10% VDD and a wide temperature range from 248K to 373K. Furthermore, this RRAM PUF is compatible with in-memory computing (IMC), and they can be implemented using the same RRAM crossbar array. Jiang Li 0012, Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Shahar Kvatinsky |
DATE | 3 |
| 2024 | Joint User Association and Power Control for Cell-Free Massive MIMOabstractThis work proposes novel approaches that jointly design user equipment (UE) association and power control (PC) in a downlink user-centric cell-free massive multiple-input multiple-output (CFmMIMO) network, where each UE is only served by a set of access points (APs) for reducing the fronthaul signalling and computational complexity. In order to maximize the sum spectral efficiency (SE) of the UEs, we formulate a mixed-integer nonconvex optimization problem under constraints on the per-AP transmit power, quality-of-service rate requirements, maximum fronthaul signalling load, and maximum number of UEs served by each AP. In order to efficiently solve the formulated problem, we propose two different schemes according to the different sizes of the CFmMIMO systems. For small-scale CFmMIMO systems, we present a successive convex approximation (SCA) method to obtain a stationary solution and also develop a learning-based method (JointCFNet) to reduce the computational complexity. For large-scale CFmMIMO systems, we propose a low-complexity suboptimal algorithm using accelerated projected gradient (APG) techniques. Numerical results show that our JointCFNet can yield similar performance and significantly decrease the run time compared with the SCA algorithm in small-scale systems. The presented APG approach is confirmed to run much faster than the SCA algorithm in large-scale systems while obtaining an SE performance close to that of the SCA approach. Moreover, the median sum SE of the APG method is up to about 2.8 fold higher than that of the heuristic baseline scheme. Chongzheng Hao, Tung Thanh Vu, Hien Quoc Ngo, Minh N. Dao, Xiaoyu Dang, Chenghua Wang, Michail Matthaiou |
IEEE Internet Things J. | 6 |
| 2024 | FPAX: A Fast Prior Knowledge-Based Framework for DSE in Approximate ConfigurationsabstractCurrent artificial intelligence and data science applications typically require complex computations and massive amounts of data handling, presenting unprecedented challenges for embedded platforms. Approximate computing has emerged as the most promising design technique to address this issue, by providing a potential performance increase, while sacrificing accuracy within an acceptable range. Approximate arithmetic units require the creation of design space exploration techniques that can swiftly and automatically form an approximate configuration in fault-tolerant systems. Existing methods, however, use iterative design space sampling, resulting in a large amount of redundant computation. In this work, we propose the efficient FPAX automatic search framework which can learn from prior knowledge regarding the exploration process of known applications and use it to guide design exploration. This avoids excessive redundant computation and quickly provides an impressive approximate configuration. Compared with the Jump Search algorithm known for its efficiency, FPAX can also achieve faster convergence speed and better exploration quality. Even compared to our previous ENAP framework, it exhibits an 18x faster performance while achieving almost identical exploration quality for several commonly used fault-tolerant applications. Yuqin Dou, Chenghua Wang, Haroon Waris, Roger F. Woods, Weiqiang Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | An Efficient Ring Oscillator PUF Using Programmable Delay Units on FPGAabstractThe ring oscillator (RO) PUF can be implemented on different FPGA platforms with high uniqueness and reliability. To decrease the hardware cost of conventional RO PUFs, a new design using the programmable delay units is proposed, namely, PRO PUF. The programmable interconnect points (PIPs) of programmable delay units are used to enhance the configurability. The PUF cell of the proposed design has the ability to be efficiently programmed to an RO PUF at any stage by adjusting the propagation paths of the delay units. A significant number of responses can be generated by the proposed PRO PUF while consuming fewer hardware resources. To verify the performance, the proposed design has been implemented on Xilinx FPGAs and also simulated using a standard 40nm technology. The experimental results have shown that the proposed design achieves high uniqueness, reliability, and hardware efficiency. Moreover, the PRO PUF has been evaluated using a machine learning attack, the CMA-ES attack. The results have shown that the proposed structure is more resistant to common modeling attacks when compared to conventional RO-related PUF designs. Yijun Cui, Jiang Li 0012, Yunpeng Chen, Chenghua Wang, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2023 | Novel Intrinsic Physical Unclonable Function Design for Post-quantum CryptographyabstractThe hardware implementations of post-quantum cryptography (PQC) algorithms are vulnerable to fault injection attacks. As a hardware security primitive, the intrinsic physical unclonable function (PUF) is a possible countermeasure for these attacks with low resource overheads. In this work, a novel intrinsic PUF, frequency adjustable software PUF (FAS-PUF), is proposed to provide a device identification for PQC chips. The FAS-PUF is based on an inherent timing logic in the ring-learning with error (R-LWE) decryption circuit of PQC chips. The FAS-PUF uses a$256^{\ast}13^{\ast} 3$-bit input ciphertext of the decryption circuit as a challenge, and uses a 256-bit decryption output as a response with an adjustable overclocking. Since the entropy of the FAS-PUF utilises the manifested timing errors caused by the overclocking, the FAS-PUF does not need to modify the existing hardware circuits, i.e. preserves the original circuit functions, which significantly reduces hardware resource consumption and power overhead. Meanwhile, to mitigate the affection of circuits' metastablities to PUF's stability under overclocking, a dynamic clock frequency selection method is used to determine the optimal frequency point for generating PUF responses. The proposed FAS-PUF is also a Strong PUF design with a significant number of Challenge/Response Pairs (CRPs) provided. The proposed design is implemented on Xilinx Basys3 FPGAs. The experimental results show that the FAS-PUF has a good uniqueness, uniformity and stability compared with other intrinsic PUFs. Yijun Cui, Chongyan Gu, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2023 | Graph contextualized self-attention network for software service sequential recommendation
Zixuan Fu, Chenghua Wang |
Future Gener. Comput. Syst. | 2 |
| 2023 | Probability density function based data augmentation for deep neural network automatic modulation classification with limited training dataabstractAbstract Deep neural networks (DNN) based automatic modulation classification (AMC) has achieved high accuracy performance. However, DNNs are data‐hungry models, and training such a model requires a large volume of data. Insufficient training data will cause DNN models to experience overfitting and severe performance degradation. In practical AMC tasks, training the deep model with sufficient data is challenging due to the costly data collection. To this end, a novel probability density function (PDF) based data augmentation scheme and a method to determine the required minimum sampling size for data enlargement is proposed. Compared with the known image‐based augmentation scheme, the proposed waveform‐based PDF technique has low complexity and is easy to implement. Experimental results show that the required size of the training dataset is one order of magnitude smaller than the sufficient dataset in the additive white Gaussian noise channel, and effective recognition can be achieved using around 60% of the total examples under the Rayleigh channel. Moreover, the presented scheme can expand training data under frequency and phase offsets. Chongzheng Hao, Xiaoyu Dang, Xiangbin Yu 0001, Sai Li 0004, Chenghua Wang |
IET Commun. | 5 |
| 2023 | ENAP: An Efficient Number-Aware Pruning Framework for Design Space Exploration of Approximate ConfigurationsabstractApproximate computing has emerged as a new computing architecture paradigm that trades off necessary numerical accuracy for performance. Various approximation operation units such as adders and multipliers have been created and provide the basis for improving system efficiency, but it is clear, that a design space exploration (DSE) is needed if improved performance is to be systematically achieved. The challenge is to determine a suitable configuration among approximation units with different error characteristics to ensure a minimization of resources while not exceeding user-defined error constraints. In this paper, we propose the efficient number-aware pruning (ENAP) technique that can compress the search space size. Using common fault-tolerant applications, we demonstrate a compression rate up to 0.0008%, meaning that 99.9992% of invalid designs can remain unsearched. An improved genetic algorithm (GA) is subsequently proposed to improve ENAP, allowing the creation of the optimal configuration in only 2 to 3 iterations, thereby greatly improving search efficiency compared to the initial 9 iterations. We integrate these two approaches into the proposed framework, demonstrating how we can achieve better exploration results compared to state-of-the-artwork. Yuqin Dou, Chenghua Wang, Roger F. Woods, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Design of High Hardware Efficiency Approximate Floating-Point FFT ProcessorabstractThe Fast Fourier Transformation (FFT), as a high-efficiency algorithm of the Discrete Fourier Transform (DFT), is widely used in Digital Signal Processing (DSP), wireless communication systems, spectrum analysis, and image processing. Approximate computing has shown effectiveness and feasibility to enhance the hardware efficiency of these applications. However, most approximate units in previous works are designed case by case, which has low efficiency and is difficult to find the optimal design. In this paper, a top-down design strategy for approximate floating-point (FP) FFT is proposed, which includes a mantissa bit-width adjustment algorithm and a step-by-step multiplier approximation algorithm. With the mantissa bit-width adjustment algorithm, the approximate 64 FP FFT achieved 50% area reduction and 70% power-delay product (PDP) reduction compared to the exact design with a 60dB Signal Noise Ratio (SNR) requirement, which is also at least 52% and 33% better than the previous approximate FP FFT. After using the step-by-step multiplier approximation algorithm, the approximate mantissa multiplier with an 8-bit fractional part reduced the area and PDP by 81.15% and 93.70%, respectively. The feasibility of the proposed approximate FFT design is verified in the channel estimation module of a wireless communication system, spectrum analysis, and image processing system. Chenggang Yan 0002, Jipeng Ge, Chenghua Wang, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | PAxC: A Probabilistic-oriented Approximate Computing Methodology for ANNsabstractIn spite of the rapidly increasing number of approximate designs in circuit logic stack for Artificial Neural Networks (ANNs) learning. A principled and systematic approximate hardware incorporating domain knowledge is still lacking. As the layer of ANN becomes deeper, the errors introduced by approximate hardware will be accumulated quickly, which can result in unexpected results. In this paper, we propose a probabilistic-oriented approximate computing (PAxC) methodology based on the notion of approximate probability to overcome the conceptual and computational difficulties inherent to probabilistic ANN learning. The PAxC makes use of minimum likelihood error in both circuit and application level to maintain the aggressive approximate datapaths to boost the benefits from the tradeoff between accuracy and energy. Compared with a baseline design, the proposed method significantly reduces the power-delay product (PDP) with a negligible accuracy loss. Simulation and a case study of image processing validate the effectiveness of the proposed methodology. Chenghua Wang, Ke Chen 0018, Weiqiang Liu 0001 |
DATE | 2 |
| 2022 | Horizontal Correlation Analysis without Precise Location on Schoolbook Polynomial Multiplication of Lattice-based CryptosystemabstractMost cryptographic systems are secure in theory; however, the implementation of cryptographic system on embedded devices can be attacked by analyzing the power consumption of specific operation to reveal the key. The classic vertical correlation power analysis (CPA) attack requires a large number of power traces for analysis. Using transient secret-key scheme significantly weakens such an attack as insufficient data could be obtained. On the other hand, the horizontal CPA requires at least a single power trace and can make full use of multiple intermediate values to analyze the correlation of power consumption. In this work, we devised a horizontal CPA attack on schoolbook polynomial multiplication of hardware-implemented lattice-based cryptosystem without precise location. The accuracy of correctly recovering any one sub secret-key using only a single trace is 99.90%, and the accuracy of correctly recovering the secret-key is 76.41%. The powerful attack capability of horizontal CPA exposes the vulnerability of unprotected schoolbook polynomial multiplication against the attack of side-channel analysis (SCA). Chuanchao Lu, Yijun Cui, Dur-e-Shahwar Kundi, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2022 | AxRLWE: A Multilevel Approximate Ring-LWE Co-Processor for Lightweight IoT ApplicationsabstractThis work presents a multilevel approximation exploration undertaken on the Ring-Learning-with-Errors (R-LWE)-based public-key cryptographic (PKC) schemes that belong to quantum-resilient cryptography algorithms. Among the various quantum-resilient cryptography schemes proposed in the currently running NIST’s post-quantum cryptography (PQC) standardization plan, the lattice-based learning-with-error (LWE) schemes have emerged as the most viable and preferred class for the Internet of Things (IoT) applications due to their compact area and memory footprint compared to other alternatives. However, compared to the classical schemes used today, R-LWE is much harder a challenge to fit on embedded IoT (end-node) devices, due to their stricter resource constraints (lower area, memory, and energy budgets) as well as their limited computational capabilities. To the best of our knowledge, this is the first endeavor exploring the inherent approximate nature of the LWE problem to undertake a multilevel approximate R-LWE (AxRLWE) architecture with respective security estimates opt for lightweight IoT devices. Undertaking AxRLWE on field-programmable gate arrays (FPGAs), we benchmarked a 64% area reduction cost compared to earlier accurate R-LWE designs at the cost of reduced quantum security. For the application-specific integrated circuits (ASICs) with 45-nm CMOS technology, AxRLWE was benchmarked to fit well within the same area budget of a lightweight ECC processor and consume a third of energy compared to special class of R-Binary LWE (R-BLWE) designs being proposed for an IoT, with a better security level. Dur-e-Shahwar Kundi, Ayesha Khalid, Song Bian 0001, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Internet Things J. | 4 |
| 2022 | A Generic Dynamic Responding Mechanism and Secure Authentication Protocol for Strong PUFsabstractAs a lightweight hardware security primitive, physical unclonable functions (PUFs) can provide reliable identity authentication for devices of Internet of Things (IoTs) with limited resources. However, the delay-based PUF structures in authentication protocols have static responding behaviors, which make them vulnerable to modeling attacks. To address this issue, many complex PUF designs have been designed to increase the nonlinearity of their models. However, most of them can still be broken by modeling-based machine learning (ML) attacks. In this article, a dynamic responding mechanism for PUF designs to generate dynamic responses is proposed. Different from the concept of logically reconfigurable PUFs, the proposed mechanism does not rely on external inputs to provide reconfiguration signals. And different from the conventional PUF authentication protocols that use large-size linear feedback shift register (LFSR) to extend the master challenge, the proposed scheme uses internally generated dynamic signals to obfuscate the master challenge to generate multiple subchallenges. These subchallenges are then input to the underlying strong PUF to generate multibit dynamic responses. It can prevent an attacker from obtaining valid challenge-response pairs (CRPs) for the underlying PUF. A security authentication protocol is also proposed, the special authentication bit-string design can resist both conventional ML attacks and the latest covariance matrix adaptation evolution strategies (CMA-ES) variant. Yale Wang, Chenghua Wang, Chongyan Gu, Yijun Cui, Máire O'Neill, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | A 10-b 500MS/s Partially Loop-Unrolled SAR ADC with a Comparator Offset Calibration TechniqueabstractThis paper presents a 10-b 500MS/s successiveapproximation-register (SAR) analog-to-digital converter (ADC) designed using a 40nm CMOS process. The first 6-bit coarse conversion is completed by a high speed loop-unrolled architecture, while the succeeding 5 bits are obtained by a traditional SAR structure. A foreground calibration is employed to correct the offsets in the six comparators of the coarse converter, while the residual errors due to process-voltage-temperature (PVT) variations are covered by 1-bit redundancy. A background offset calibration technique based on alternate comparators is proposed, which tracks PVT variations while eliminating a dedicated calibration phase. The spurious-free-dynamic-range (SFDR) and the signal-to-noise-and-distortion-ratio (SNDR) can achieve 60.30dB and 68.95dBc, respectively. The power consumption of the whole system is 4.164mW under 1.1V supply voltage, thereby obtaining a figure of merit (FoM) of 9.87fJ/conv.-step. Jie Sun 0020, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2021 | Towards CRYSTALS-Kyber: A M-LWE Cryptoprocessor with Area-Time Trade-OffabstractCRYSTALS-Kyber is a quantum-resistant and promising lattice-based cryptography (LBC) in the finalists of the third round post-quantum cryptography (PQC) standardization, which is based on the hardness of Module-Learning with Errors (M-LWE). The variadic parameters make M-LWE obtain a more flexible security-performance trade-off than Ring-LWE. In this paper, we propose a M-LWE cryptoprocessor targeting CRYSTALS-Kyber with area-time trade-off for the first time. This balanced design includes a fast and low-cost Binomial Sampler and vector-polynomials multiplication structure based on pipelined decimation-in-frequency (DIF) based Number Theoretic Transform (NTT) technique. The M-LWE cryptoprocessor achieve 27,708 encryption operations per second using only 690 slices and 106,716 decryption operations per second using only 571 slices. Our proposed design achieved the lowest area-time product (ATP) with at least 2 χ performance improvement than the state-of-the-art LBC designs with a similar security level and complexity of polynomials. Kan Yao, Dur-e-Shahwar Kundi, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2021 | A Dynamic Highly Reliable SRAM-Based PUF Retaining Memory FunctionabstractIn this paper, a highly reliable SRAM based Physical Unclonable Function (PUF), which retains the memory function is proposed. The mismatch of NMOS is extracted during discharge process and amplified by the cross-coupled inverter to generate a response. At the beginning of the discharge process, the NMOSs are biased at sub-threshold region, which can improve the reliability and stability. The proposed PUF is designed in a 40nm CMOS process and each bit cell only consumes 4.98 μm2(3112F2). Post simulation shows that the bit error rate (BER) deterioration is 0.96% per 0.1V, 0.36% per 10° C with temperature variations from -40° C to 80° C and supply voltage variations from 0.9V to 1.3V. It achieves 1.8% native instability through the simulation. Meanwhile, the proposed PUF can retain memory function after a response is generated. Chenghua Wang, Chenggang Yan 0002, Yijun Cui, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2021 | An Energy Efficient Accelerator for Bidirectional Recurrent Neural Networks (BiRNNs) Using Hybrid-Iterative Compression With Error SensitivityabstractRecurrent Neural Networks (RNNs) have been widely used in many sequential applications, such as machine translation, speech recognition and sentiment analysis. Long Term Short Term Memory (LSTM) and Gated Recurrent Unit (GRU) are widely used variants of RNN due to their effectiveness in overcoming gradient vanishing and exploding problems; however, compared to conventional RNN, their massive storage and computation requirements hinder their application. In addition, the recurrent structure of RNNs makes them prone to accumulate errors, resulting in a severe loss of accuracy. In this work, we propose a hybrid-iterative compression (HIC) algorithm for LSTM/GRU. By exploiting the error sensitivity of RNN, the gating units are divided into error-sensitive and error-insensitive groups, that are compressed using different algorithms. By using this approach, a 37.1×/32.3× compression ratio is achieved with negligible accuracy loss for LSTM/GRU. Further, an energy efficient accelerator for bidirectional RNNs is proposed. In this accelerator, the data flow of the matrix operation unit based on the block structure matrix (MOU-S) is improved through rearranging weights; the utilization of BRAM is improved through a fine-grained parallelism configuration of matrix-vector multiplications (MVMs). Meanwhile, the timing matching strategy alleviates the load-imbalance problem between MOU-S and the matrix operation unit based on top- k pruning (MOU-P). When running at 200MHz on Xilinx ADM-PCIE-7V3 FPGA, the proposed design achieves an improvement in energy efficiency in a range of 5%-237% for LSTM networks, and an improvement of 58% for GRU networks compared with state-of-the-art designs. Guocai Nan, Zhengkuan Wang, Chenghua Wang, Bi Wu 0002, Zhican Wang, Weiqiang Liu 0001, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Design and Analysis of Energy-Efficient Dynamic Range Approximate Logarithmic Multipliers for Machine LearningabstractApproximate computing provides an emerging approach to design high performance and low power arithmetic circuits. The logarithmic multiplier (LM) converts multiplication into addition and has inherent approximate characteristics. In this article, dynamic range approximate LMs (DR-ALMs) for machine learning applications are proposed; they use Mitchell’s approximation and a dynamic range operand truncation scheme. The worst case (absolute and relative) errors for the proposed DR-ALMs are analyzed. The accuracy and the hardware overhead of these designs are provided to select the best approximate scheme according to different metrics. The proposed DR-ALMs are compared with the conventional LM with exact operands and previous approximate multipliers; the results show that the power-delay product (PDP) of the best proposed DR-ALM (DR-ALM-6) are decreased by up to 54.07 percent with the mean relative error distance (MRED) decreasing by 21.30 percent compared with 16-bit conventional design. Case studies for three machine learning applications show the viability of the proposed DR-ALMs. Compared with the exact multiplier and its conventional counterpart, the back-propagation classifier with DR-ALMs with a truncation length larger than 4 has a similar classification result for the three datasets; the K-means clustering application with all DR-ALMs has a similar clustering result for four datasets; and the handwritten digit recognition application with DR-ALM-5 or DR-ALM-6 for LeNet-5 achieves similar or even slightly higher recognition rate. Peipei Yin, Chenghua Wang, Haroon Waris, Weiqiang Liu 0001, Yinhe Han 0001, Fabrizio Lombardi |
IEEE Trans. Sustain. Comput. | 2 |
| 2020 | Security Analysis of Hardware Trojans on Approximate CircuitsabstractApproximate computing, for error-tolerant applications, provides trade-offs for computations to achieve improved speed and power performance. Approximate circuits, in particular approximate arithmetic circuits, directly affect the performance of a computing system. Hence, approximate circuit designs have been extensively studied. However, security issues of approximate circuits have been ignored. Moreover, hardware Trojans have been found in fabricated chips in manufacturing industry chains by untrusted foundries. Hardware Trojans could affect the functionality of approximate circuits under very rare circumstances with inconsiderable footprints. In this paper, hardware Trojan insertion methods based on signal transition probability are utilized to investigate and evaluate the security threats in approximate circuits. A approximate low-partor-adder (LOA) adder is utilized as an example and analyzed in the paper. The evaluation results show that with the increase of the number of approximation modules, the approximate LOA adder is more possible to be inserted hardware Trojans than the exact LOA adder. Yuqin Dou, Shichao Yu, Chongyan Gu, Máire O'Neill, Chenghua Wang, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2020 | Programmable Ring Oscillator PUF Based on Switch MatrixabstractConfigurable ring oscillator (CRO) physical unclonable functions (PUFs) which can improve the uniqueness and reliability of conventional RO PUFs have been widely studied. Especially, the multiplier, XOR gate and tristate inverter based CRO PUFs can improve the uniqueness and reliability. However the efficiency is remain at the same level when compared with the conventional RO PUFs. In this paper, a programmable RO PUF (PRO PUF), which can be programmed to change the structure of a typical RO PUF, is proposed. The proposed PRO PUF design is implemented based on the switch matrix of an FPGA and can be programmed as a chained RO PUF or a random looped RO PUF. The proposed PRO PUF is implemented on Xilinx Spartan 6 FPGAs. Experimental results demonstrate that the proposed PRO PUF design has good uniqueness and reliability metrics as well as a high hardware efficiency. Yijun Cui, Yunpeng Chen, Chenghua Wang, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2020 | AxMM: Area and Power Efficient Approximate Modular Multiplier for R-LWE CryptosystemabstractAmongst various Post-Quantum Cryptographic (PQC) schemes, Lattice-Based Cryptography (LBC) stands out as the most viable substitute to the classical cryptographic schemes due to its efficiency, versatility and solid foundations on hard mathematical problems. Ring Learning With Errors (R-LWE) is a Public Key Encryption (PKE) scheme of LBC, in which the modular polynomial multiplication in a ring is the main bottleneck in the realization of a practical resource-constraint design for the embedded IoT devices. This work explores novel Approximate Computing (AC) technique for the design of area/power efficient modular multiplier (so called AxMM) for R-LWE, exploiting the inherent approximate structure of the scheme. The proposed AxMM on 45nm ASIC library achieved an area and power reduction of 36% and 23%, respectively, along with a speed increase of 1.34× as compared to state-of-art smallest exact R-LWE modular multiplier. Dur-e-Shahwar Kundi, Song Bian 0001, Ayesha Khalid, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2020 | DC-LSTM: Deep Compressed LSTM with Low Bit-Width and Structured MatricesabstractLong Short-Term Memory (LSTM) has been widely adopted in many sequential applications, such as language model and speech recognition. LSTM usually incurs in a large memory requirement and high computational complexity. Therefore, LSTM has a limited applicability to embedded and mobile systems. In LSTM, a large number of operations and high storage are required for matrix-vector multiplication (MV). In this paper, we present a software and hardware co-design scheme for efficiently compressing MVs. By utilizing a structured matrix, quantization and selective top-k pruning, memory requirements are substantially reduced while only incurring in a negligible accuracy loss. Then, a block-parallel hardware architecture is proposed for the compressed LSTM. As requiring less multiplication operations and storage resources, the proposed architecture achieves the very good compression ratio. The proposed architecture is implemented on the Xilinx VCU118 and KC705 platforms. Experimental results show that the proposed design uses less DSP and BRAM resources. Guocai Nan, Chenghua Wang, Weiqiang Liu 0001, Fabrizio Lombardi |
ISCAS | 2 |
| 2019 | Theoretical Analysis of Delay-Based PUFs and Design Strategies for ImprovementabstractDelay-based physical unclonable function (PUF) designs use the random delay differences in circuit transmission to extract response. In the existing PUF designs, there are few studies on investigating the link between process variation and PUF performance. The experimental data can reflect the performance of the new design to a certain extent, but lack of theoretical analysis to provide thorough information. In this paper, a theoretical model for delay-based PUF designs is proposed. An analysis of the delay-based PUF improvements by existing design strategies is also investigated. Moreover, a guidance to develop and improve future delay-based PUF designs using the proposed theoretical model is also given in this paper. Yale Wang, Chenghua Wang, Chongyan Gu, Yijun Cui, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2019 | Design and Analysis of Approximate Redundant Binary MultipliersabstractAs technology scaling is reaching its limits, new approaches have been proposed for computional efficiency. Approximate computing is a promising technique for high performance and low power circuits as used in error-tolerant applications. Among approximate circuits, approximate arithmetic designs have attracted significant research interest. In this paper, the design of approximate redundant binary (RB) multipliers is studied. Two approximate Booth encoders and two RB 4:2 compressors based on RB (full and half) adders are proposed for the RB multipliers. The approximate design of the RB-Normal Binary (NB) converter in the RB multiplier is also studied by considering the error characteristics of both the approximate Booth encoders and the RB compressors. Both approximate and exact regular partial product arrays are used in the approximate RB multipliers to meet different accuracy requirements. Error analysis and hardware simulation results are provided. The proposed approximate RB multipliers are compared with previous approximate Booth multipliers; the results show that the approximate RB multipliers are better than approximate NB Booth multipliers especially when the word size is large. Case studies of error-resilient applications are also presented to show the validity of the proposed designs. Weiqiang Liu 0001, Tian Cao 0005, Peipei Yin, Yuying Zhu 0003, Chenghua Wang, Earl E. Swartzlander Jr., Fabrizio Lombardi |
IEEE Trans. Computers | 5 |
| 2019 | XOR-Based Low-Cost Reconfigurable PUFs for IoT SecurityabstractWith the rapid development of the Internet of Things (IoT), security has attracted considerable interest. Conventional security solutions that have been proposed for the Internet based on classical cryptography cannot be applied to IoT nodes as they are typically resource-constrained. A physical unclonable function (PUF) is a hardware-based security primitive and can be used to generate a key online or uniquely identify an integrated circuit (IC) by extracting its internal random differences using so-called challenge-response pairs (CRPs). It is regarded as a promising low-cost solution for IoT security. A logic reconfigurable PUF (RPUF) is highly efficient in terms of hardware cost. This article first presents a new classification for RPUFs, namely circuit-based RPUF (C-RPUF) and algorithm-based RPUF (A-RPUF); two Exclusive OR (XOR)-based RPUF circuits (an XOR-based reconfigurable bistable ring PUF (XRBR PUF) and an XOR-based reconfigurable ring oscillator PUF (XRRO PUF)) are proposed. Both the XRBR and XRRO PUFs are implemented on Xilinx Spartan-6 field-programmable gate arrays (FPGAs). The implementation results are compared with previous PUF designs and show good uniqueness and reliability. Compared to conventional PUF designs, the most significant advantage of the proposed designs is that they are highly efficient in terms of hardware cost. Moreover, the XRRO PUF is the most efficient design when compared with previous RPUFs. Also, both the proposed XRRO and XRBR PUFs require only 12.5% of the hardware resources of previous bitstable ring PUFs and reconfigurable RO PUFs, respectively, to generate a 1-bit response. This confirms that the proposed XRBR and XRRO PUFs are very efficient designs with good uniqueness and reliability. Weiqiang Liu 0001, Lei Zhang 0089, Zhengran Zhang, Chongyan Gu, Chenghua Wang, Máire O'Neill, Fabrizio Lombardi |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2018 | Combining Restoring Array and Logarithmic Dividers into an Approximate Hybrid DesignabstractThis paper proposes a new design of an approximate hybrid divider (AXHD), which combines the restoring array and the logarithmic dividers to achieve an excellent tradeoff between accuracy and hardware performance. Exact restoring divider cells (EXDCrs) are used to generate the MSBs of the quotient for attaining a high accuracy; the other quotient digits are processed by a logarithmic divider as inexact scheme to improve figures of merit such as power consumption, area and delay. The proposed AXHD is evaluated and analyzed using error and hardware metrics. The proposed design is also compared with the exact restoring divider (EXDr) and previous approximate restoring dividers (AXDrs). The results show that the proposed design achieves very good performance in terms of accuracy and hardware; case studies for image processing also show the validity of the proposed designs. Weiqiang Liu 0001, Jing Li 0117, Chenghua Wang, Paolo Montuschi, Fabrizio Lombardi |
ARITH | 4 |
| 2018 | A machine learning attack resistant multi-PUF design on FPGAabstractCurrent approaches for building physical unclonable function (PUF) designs resistant to machine learning attacks often suffer from large resource overhead and are typically difficult to implement on field programmable gate arrays (FPGAs). In this paper we propose a new arbiter-based multi-PUF (MPUF) design that utilises a Weak PUF to obfuscate the challenges to a Strong PUF and is harder to model than the conventional arbiter PUF using machine learning attacks. The proposed PUF design shows a greater resistance to attacks, which have been successfully applied to other Arbiter PUFs. A mathematical model is presented to analyse the complexity and obfuscation properties of the proposed PUF design. Moreover, we show that it is feasible to implement the proposed MPUF design on a Xilinx Artix-7 FPGA, and that it achieves a good uniqueness result of 40.60 % and uniformity of 37.03 %, which significantly improves over previous work into multi-PUF designs. Qingqing Ma, Chongyan Gu, Neil Hanley, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill |
ASP-DAC | 4 |
| 2018 | Design of Dynamic Range Approximate Logarithmic MultipliersabstractApproximate computing is an emerging approach for designing high performance and low power arithmetic circuits. The logarithmic multiplier (LM) converts multiplication into addition and has inherent approximate characteristics. A method combining the Mitchell's approximation and a dynamic range operand truncation scheme is proposed in this paper to design non-iterative and iterative approximate LMs. The accuracy and the circuit requirements of these designs are assessed to select the best approximate scheme according to different metrics. Compared with conventional non-iterative and iterative 16-bit LMs with exact operands, the normalized mean error distance (NMED) of the best proposed approximate non-iterative and iterative LMs is decreased up to 24.1% and 18.5%, respectively, while the power-delay product (PDP) is decreased up to 51.7% and 45.3%, respectively. Case studies for two error-tolerant applications show the validity of the proposed approximate LMs. Peipei Yin, Chenghua Wang, Weiqiang Liu 0001, Fabrizio Lombardi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Design of Approximate FFT with Bit-width Selection AlgorithmsabstractThis paper presents the approximate designs of Fast Fourier Transformation (FFT) circuit. The tradeoff between accuracy and hardware performance is achieved by using bit-width selection for each stage. The error rate can be tuned with bit-width selection. We proposed two algorithms for bit-width selection under certain error restriction. The first algorithm is targeting an approximate FFT design with low hardware cost. While the second algorithm is proposed to achieve high performance. Both of proposed algorithms allow the designer to tradeoff hardware performance and computation accuracy in each stage. The proposed two designs are implemented on FPGA. The results show that the approximate FFT design using the first algorithm can reduce hardware resource consumption up to 30.2%. The second algorithm can increases the performance of the approximate FFT design up to 24.0%, while it also saves 25.2% resource consumption. Qicong Liao, Weiqiang Liu 0001, Fei Qiao, Chenghua Wang, Fabrizio Lombardi |
ISCAS | 4 |
| 2017 | XOR gate based low-cost configurable RO PUFabstractA Physical Unclonable Function (PUF) is often used to uniquely identify an integrated circuit by extracting its internal random differences using so-called Challenge Response Pairs (CRPs). As CRPs include unique information about the underlying hardware variations, PUF design is a promising approach to provide authentication and IP-protection capabilities. In this paper, an XOR-gate-based configurable Ring Oscillator (RO) PUF (denoted as XCRO PUF) is presented. This XCRO PUF can generate more CRPs compared with state-of-the-art PUF designs by using the same number of configurable logic blocks (CLBs) in an FPGA implementation. This design is implemented in the Xilinx Spartan-6 XC6SLX9 FPGAs with fixed locations for the XCROs (placed within a ring to improve its uniqueness). The XCRO PUF shows better uniqueness and reliability than other PUF designs. Moreover, a XCRO PUF consumes only 12.5% of the hardware resources to generate a 1-bit response compared with other CRO PUFs implemented in FPGA. Lei Zhang 0089, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill, Fabrizio Lombardi |
ISCAS | 2 |
| 2017 | Design of Approximate Radix-4 Booth Multipliers for Error-Tolerant ComputingabstractApproximate computing is an attractive design methodology to achieve low power, high performance (low delay) and reduced circuit complexity by relaxing the requirement of accuracy. In this paper, approximate Booth multipliers are designed based on approximate radix-4 modified Booth encoding (MBE) algorithms and a regular partial product array that employs an approximate Wallace tree. Two approximate Booth encoders are proposed and analyzed for error-tolerant computing. The error characteristics are analyzed with respect to the so-called approximation factor that is related to the inexact bit width of the Booth multipliers. Simulation results at 45 nm feature size in CMOS for delay, area and power consumption are also provided. The results show that the proposed 16-bit approximate radix-4 Booth multipliers with approximate factors of 12 and 14 are more accurate than existing approximate Booth multipliers with moderate power consumption. The proposed R4ABM2 multiplier with an approximation factor of 14 is the most efficient design when considering both power-delay product and the error metric NMED. Case studies for image processing show the validity of the proposed approximate radix-4 Booth multipliers. Weiqiang Liu 0001, Liangyu Qian, Chenghua Wang, Honglan Jiang, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 3 |
| 2016 | Live demonstration: An automatic evaluation platform for physical unclonable function testabstractPUF is a security primitive that exploits the fact that no two ICs are exactly the same. To verify a new PUF design, several metrics including uniqueness, reliability, and randomness must be evaluated, which requires various resources and a long set-up time. In this live demonstration, we have developed an automatically evaluation platform for the PUF design. To the authors' best knowledge, this is the first automatic evaluation platform for the PUF test. The evaluation platform can be used for both FPGA and ASCI PUF testing. Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill |
ISCAS | 2 |
| 2016 | Low-cost configurable ring oscillator PUF with improved uniquenessabstractThe physical unclonable function (PUF) produces die-unique responses and is regarded as an emerging security primitive that can be used for authentication of devices. The complexity of a conventional PUF design based on a ring oscillator (RO) is rather high, so limiting its use in many applications. The configurable ring oscillator (CRO) PUF has been advocated as a possible solution to this issue. In this paper, a low hardware complexity CRO PUF design with an enhanced capability to generate a large number of bit responses is proposed; only an inverter and a multiplexer are used in each delay unit. The responses are generated by considering the variation due to fabrication of the logic gates and wires in the CROs. A novel comparison strategy is proposed for the generation of the responses. The proposed PUF design is implemented on Xilinx Spartan-6 FPGAs. These results show that the proposed CRO PUF design has good uniqueness; moreover, it is also robust in its operation for the temperature range of -25°C~85°C. Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Máire O'Neill, Fabrizio Lombardi |
ISCAS | 2 |
| 2016 | Design and evaluation of an approximate Wallace-Booth multiplierabstractApproximate or inexact computing has recently attracted considerable attention due to its potential advantages with respect to high performance and low power consumption. This paper presents the design of an approximate multiplier; this approximate multiplier consists of an approximate Booth encoder, an approximate 4-2 compressor and an approximate tree structure. The approximate design is implemented and verified for 8×8, 16×16 and 32×32-bit signed multiplication schemes targeting applications in embedded systems. Simulation results at 45 nm technology are provided and discussed. Compared with an exact Wallace-Booth multiplier as well as other approximate multipliers found in the technical literature, the proposed approximate scheme achieves significant improvements in power consumption, delay and combined metrics. These results show the viability of the proposed design. Liangyu Qian, Chenghua Wang, Weiqiang Liu 0001, Fabrizio Lombardi, Jie Han 0001 |
ISCAS | 2 |
| 2016 | Design and Analysis of Inexact Floating-Point AddersabstractPower has become a key constraint in nanoscale integrated circuit design due to the increasing demands for mobile computing and higher integration density. As an emerging computational paradigm, an inexact circuit offers a promising approach to significantly reduce both dynamic and static power dissipation for error-tolerant applications. In this paper, an inexact floating-point adder is proposed by approximately designing an exponent subtractor and mantissa adder. Related operations such as normalization and rounding are also dealt with in terms of inexact computing. An upper bound error analysis for the average case is presented to guide the inexact design; it shows that the inexact floating-point adder design is dependent on the application data range. High dynamic range images are then processed using the proposed inexact floating-point adders to show the validity of the inexact design; comparison results show that the proposed inexact floating-point adders can improve the power consumption and power-delay product by 29.98 and 39.60 percent, respectively. Weiqiang Liu 0001, Linbin Chen, Chenghua Wang, Máire O'Neill, Fabrizio Lombardi |
IEEE Trans. Computers | 3 |
| 2015 | RO PUF design in FPGAs with new comparison strategiesabstractA Physical Unclonable Function (PUF) can be used to provide authentication of devices by producing die-unique responses. In PUFs based on ring oscillators (ROs), the responses are derived from the oscillation frequencies of the ROs. However, RO PUFs can be vulnerable to attack due to the frequency distribution characteristics of the RO arrays. In this paper, in order to improve the design of RO PUFs for FPGA devices, the frequencies of RO arrays implemented on a large number of FPGA chips are statistically analyzed. Three RO frequency distribution (ROFD) characteristics are observed and discussed. Based on these ROFD characteristics, two RO comparison strategies are proposed that can be used to improve the design of RO PUFs. It is found that the symmetrical RO comparison strategy has the highest entropy density. Weiqiang Liu 0001, Chenghua Wang, Yijun Cui, Máire O'Neill |
ISCAS | 3 |