Yuan Cao 0003

dblp:52/4472-3 · DBLP profile ↗
← Back
42ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0001-5227-2241ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 8 first-author · 16 since 2021Security and privacy · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Live Demonstration: A 0.787pJ/bit, 25 Mbps True Random Number Generator based on Analog Chaotic Tent Map
Ruiqi Bao, Yuan Cao 0003, Guangjun Yin
ISCAS2
2026 An Efficient Sine/Cosine Design Using Piecewise Quadratic Approximation for FOC Applications
Yuan Gao 0007, Yuan Cao 0003, Jianjun Zhuang, Rongkai Pan, Jing Tian 0004
ISCAS3
2026 Same Last-Item Confusion Unveiled: A Unified Mitigation Framework for Graph Learning in Session-Based Recommendation
abstract
Session-based recommendation (SBR), which focuses on next-item prediction for anonymous users based on short-term interaction sequences, has garnered increasing attention from researchers. While graph neural networks (GNNs) have become predominant in modeling complex item transition patterns, our empirical study reveals two critical limitations in existing GNN-based SBR methods. On the one hand, they struggle to differentiate between sessions sharing the same last item, resulting in indistinguishable session representations. On the other hand, the inherent popularity bias in session data leads to the over-recommendation of popular items. Inspired by contrastive learning techniques, this paper presents a unified mitigation framework for Same lAst-item confusion in Graph lEarning (SAGE) for SBR. In SAGE, we first obtain normalized session embeddings on constructed session graphs. We then build positive and negative samples of sessions through dual forward propagations and a novel negative sample selection strategy, followed by calculating contrastive loss. Finally, the enhanced session embeddings are utilized for prediction. Extensive experiments on two real-world datasets demonstrate that integrating SAGE with various state-of-the-art GNN-based SBR methods significantly improves their original performances.
Jinpeng Chen 0001, Jianxiang He, Yuan Cao 0003, Huan Li 0003, Zhenye Yang, Kaimin Wei, Xiongnan Jin, Senzhang Wang, Weiping Tu
WWW3
2025 Leveraging Multimodal Data and Side Users for Diffusion Cross-Domain Recommendation
abstract
Cross-domain recommendation (CDR) aims to address the persistent cold-start problem in Recommender Systems. Current CDR research concentrates on transferring cold-start users' information from the auxiliary domain to the target domain. However, these systems face two main issues: the underutilization of multimodal data, which hinders effective cross-domain alignment, and the neglect of side users who interact solely within the target domain, leading to inadequate learning of the target domain's vector space distribution. To address these issues, we propose a model leveraging Multimodal data and Side users for diffusion Cross-domain recommendation (MuSiC). We first employ a multimodal large language model to extract item multimodal features and leverage a large language model to uncover user features. Secondly, we propose the cross-domain diffusion module to learn the generation of feature vectors in the target domain. This approach involves learning feature distribution from side users and understanding the patterns in cross-domain transformation through overlapping users. Subsequently, the trained diffusion module is used to generate feature vectors for cold-start users in the target domain, enabling the completion of cross-domain recommendation tasks. Finally, our experimental evaluation of the Amazon dataset confirms that MuSiC achieves state-of-the-art performance, significantly outperforming all selected baselines. Our code is available: https://github.com/zhangf16/MuSiC.
Jinpeng Chen 0001, Huan Li 0003, Senzhang Wang, Yuan Cao 0003, Kaimin Wei, Jianxiang He, Feifei Kou, Jinqing Wang
ACM Multimedia5
2025 Hierarchical Intent-guided Optimization with Pluggable LLM-Driven Semantics for Session-based Recommendation
abstract
Session-based Recommendation (SBR) aims to predict the next item a user will likely engage with, using their interaction sequence within an anonymous session. Existing SBR models often focus only on single-session information, ignoring inter-session relationships and valuable cross-session insights. Some methods try to include inter-session data but struggle with noise and irrelevant information, reducing performance. Additionally, most models rely on item ID co-occurrence and overlook rich semantic details, limiting their ability to capture fine-grained item features. To address these challenges, we propose a novel hierarchical intent-guided optimization approach with pluggable LLM-driven semantic learning for session-based recommendations, called HIPHOP. First, we introduce a pluggable embedding module based on large language models (LLMs) to generate high-quality semantic representations, enhancing item embeddings. Second, HIPHOP utilizes graph neural networks (GNNs) to model item transition relationships and incorporates a dynamic multi-intent capturing module to address users' diverse interests within a session. Additionally, we design a hierarchical inter-session similarity learning module, guided by user intent, to capture global and local session relationships, effectively exploring users' long-term and short-term interests. To mitigate noise, an intent-guided denoising strategy is applied during inter-session learning. Finally, we enhance the model's discriminative capability by using contrastive learning to optimize session representations. Experiments on multiple datasets show that HIPHOP significantly outperforms existing methods, demonstrating its effectiveness in improving recommendation quality. Our code is available: https://github.com/hjx159/HIPHOP.
Jinpeng Chen 0001, Jianxiang He, Huan Li 0003, Senzhang Wang, Yuan Cao 0003, Kaimin Wei, Zhenye Yang, Ye Ji 0002
SIGIR5
2024 An Ising Model-Based Parallel Tempering Processing Architecture for Combinatorial Optimization
abstract
Combinatorial optimization problems (COPs) are prevalent in various domains and present formidable challenges for modern computers. Searching for the ground state of the Ising model emerges as a promising approach to solve these problems. Recent studies have proposed some annealing processing architectures based on the Ising model, aimed at accelerating the solution of COPs. However, most of them suffer from low solution accuracy and inefficient parallel processing. This article presents a novel parallel tempering processing architecture (PTPA) based on the fully-connected Ising model to address these issues. The proposed modified parallel tempering algorithm supports multi-spin concurrent updates per replica and employs an efficient multi-replica swap scheme, with fast speed and high accuracy. Furthermore, an independent pipelined spin update architecture is designed for each replica, which supports replica scalability while enabling efficient parallel processing. The PTPA prototype is implemented on FPGA with 8 replicas, each with 1,024 fully-connected spins. It supports up to 64 spins for concurrent updates per replica and operates at 200 MHz. Different concurrency strategies are considered to further improve the efficiency of solving COPs. In the test of various G-set problems, PTPA achieves 3.2× faster solution speed along with 0.27% better average cut accuracy compared to a state-of-the-art FPGA-based Ising machine.
Yang Zhang 0120, Xiangrui Wang, Gaopeng Fan, Yuan Cao 0003, Yiqiu Liu, Yongkui Yang, Enyi Yao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 CareFL: Contribution Guided Byzantine-Robust Federated Learning
abstract
Byzantine-robust federated learning (FL) endeavors to empower service providers in acquiring a precise global model, even in the presence of potentially malicious FL clients. While considerable strides have been taken in the development of robust aggregation algorithms for FL in recent years, their efficacy is confined to addressing particular forms of Byzantine attacks, and they exhibit vulnerabilities when confronted with a spectrum of attack vectors. Notably, a prevailing issue lies in the heavy reliance of these algorithms on the examination of local model gradients. It is worth noting that an attacker possesses the ability to manipulate a carefully chosen small gradient of a model within a context where there could be millions of gradients available, thereby facilitating adaptive attacks. Drawing inspiration from the foundational Shapley value methodology in game theory, we introduce an effective FL scheme namedCareFL. This scheme is designed to provide robustness against a spectrum of state-of-the-art Byzantine attacks. Unlike approaches that rely on the examination of gradients,CareFLemploys a universal metric, the loss of the local model—independent of specific gradients, to identify potentially malicious clients. Specifically, in each aggregation round, the FL server trains a reference model using a small auxiliary dataset— the auxiliary dataset can be removed with a slight defense degradation trade-off. It employs the Shapley value to assess the contribution of each client-submitted model in minimizing the global model loss. Subsequently, the server selects client models closer to the reference model in terms of Shapley values for the global model update. To reduce the computational overhead ofCareFLwhen the number of clients is relatively scaled-up, we construct its variant, namelyCareFL+ generally by grouping clients. Extensive experimentation conducted on well-established MNIST and CIFAR-10 datasets, encompassing diverse model architectures, including AlexNet, demonstrates thatCareFLconsistently achieves accuracy levels comparable to those attained under attack-free conditions when faced with five formidable attacks.CareFLand CareFL+ outperform six existing state-of-the-art Byzantine-robust FL aggregation methods, includingFLTrust, across both IID and non-IID data distribution settings.
Qihao Dong, Shengyuan Yang, Zhiyang Dai, Yansong Gao 0001, Shang Wang 0004, Yuan Cao 0003, Anmin Fu, Willy Susilo
IEEE Trans. Inf. Forensics Secur.6
2023 A Template Attack on Reduction Without Reference Device on Kyber
abstract
In July 2022, the National Institute of Standards and Technology (NIST) announced its selection of four algorithms for post-quantum cryptography standardization in advance. Among these algorithms, Kyber was chosen as the only key encapsulation mechanism (KEM). In the Kyber KEM, the modular reduction function is utilized in numerous areas. We have discovered that by modeling controllable modular reduction functions, unknown modular reduction functions can be targeted. And attacks can then be constructed. Henceforth, profiling can be mounted on the target device. In this paper, we present a machine-learning-based key recovery attack on Kyber, without needing a reference device. We have effectively attacked the modular reduction function. Furthermore, this vulnerability that enables the reuse of the same function could be utilized in other attacks.
Yipei Yang, Junying Huang, Zongyue Wang, Jing Ye 0001, Junfeng Fan, Huawei Li 0001, Xiaowei Li 0001, Yuan Cao 0003
ATS10
2023 Online Reliability Evaluation Design: Select Reliable CRPs for Arbiter PUF and Its Variants
abstract
Physical Unclonable Function (PUF) is a hardware security primitive with broad application prospects. Variants of the arbiter PUF have been proposed to resist modeling attacks. However, their low reliability issue limits their applications. To solve the low reliability issue, this paper proposes an Online Reliability Evaluation (ORE) design for the arbiter PUF and its variants. Moreover, a corresponding machine learning method to select reliable Challenge Response Pairs (CRPs) for applications is proposed. Based on the ORE design, a small number of CRPs and their reliability levels are collected during the enrollment phase. Then they are trained to build reliability models for predicting the responses and reliability levels of other challenges. Since the ORE design does not change the security structures of the arbiter PUF and its variants, the resistance to modeling attacks of PUF designs equipped with it is maintained. Compared to the previous work that tests 100,000 times per CRP, our design is time-saving in the enrollment phase since each CRP is only tested three times for training reliability models. The proposed design is implemented under the 40nm process. Experimental results on real chips show that all the CRPs selected by our reliability models are indeed reliable for applications, verifying the effectiveness of our method.
Chaofang Ma, Jianan Mu, Jing Ye 0001, Yuan Cao 0003, Huawei Li 0001, Xiaowei Li 0001
ETS5
2023 Mandari: Multi-Modal Temporal Knowledge Graph-aware Sub-graph Embedding for Next-POI Recommendation
abstract
Next-POI recommendation aims to explore from user check-in sequence to predict the next possible location to be visited. Existing methods are often difficult to model the implicit association of multi-modal data with user choices. Moreover, traditional methods struggle to fully explore the variation of user preferences at variable time intervals. To tackle these limitations, we propose a Multi-Modal Temporal Knowledge Graph-aware Sub-graph Embedding approach (Mandari). We first construct a novel Multi-Modal Temporal Knowledge Graph. Based on the proposed knowledge graph, we integrate multi-modal information and leverage the graph attention network to calculate sub-graph prediction probability. Next, we implement a temporal knowledge mining method to model the segmentation and periodicity of user check-in and obtain temporal prediction probability. Finally, we fuse temporal prediction probability with the previous sub-graph prediction probability to obtain the final result. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art methods.
Xiuyun Li, Yuan Cao 0003, Xiongnan Jin, Jinpeng Chen 0001
ICME3
2023 Chosen ciphertext correlation power analysis on Kyber
Yipei Yang, Zongyue Wang, Jing Ye 0001, Junfeng Fan, Huawei Li 0001, Xiaowei Li 0001, Yuan Cao 0003
Integr.8
2023 Scalable and Conflict-Free NTT Hardware Accelerator Design: Methodology, Proof, and Implementation
abstract
Number theoretic transform (NTT) is useful for the acceleration of polynomial multiplication, which is the main performance bottleneck in the next-generation cryptographic schemes. Different NTT-based cryptographic algorithms have different security settings. The diverse application scenarios introduce different cost-performance tradeoffs and hardware constraints. Motivated by the emerging demand for more versatile NTT hardware accelerators, we propose a new design methodology that can generate area-efficient and high-performance NTT accelerators for any length and modulus of NTT polynomials and single processing element (PE) or PE array with a varying number of layers. The proposed NTT accelerator architecture pivots on a conflict-free memory access pattern for adaptation to different combinations of security and PE array configuration parameters. The proposed memory access pattern is formally proved to be conflict-free for any parametric configurations. The criterion for read-after-write conflict without pipeline stall is also established. Our proposed design methodology can produce NTT accelerators with single PE or multilayer PE array for different polynomial size and modulus, with hardware area and computational efficiency comparable to accelerators customized for a fixed set of parameters. Our proposed methodology produces parameterized accelerator with higher scalability than the existing parameterized accelerator design. On average, the accelerators generated by our proposed method are 71.4% more area-time efficient. Up to 30.7% area-time reduction over the most area-time efficient state-of-the-art scalable NTT accelerator can be achieved for the same security parameters.
Jianan Mu, Wen Wang 0007, Yizhong Hu, Chip-Hong Chang, Junfeng Fan, Jing Ye 0001, Yuan Cao 0003, Huawei Li 0001, Xiaowei Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2023 A New Reconfigurable True Random Number Generator and Physical Unclonable Function Unified Chip With On-Chip Auto-Calibration
abstract
True random number generator (TRNG) and physical unclonable function (PUF) have been extensively used to secure low-cost Internet of Things (IoT) endpoints. In this paper, a lightweight reconfigurable TRNG and PUF unified design for custom chip implementation is proposed. The reconfigurable structure consists of a pair of ring oscillators (ROs) with interposed multi-way switches for RO length reconfiguration and shared counters for on-chip calibration. Jitter noise of ROs and metastability of arbiter are harmonized for TRNG operation, while process variations of ROs are extracted for PUF operation. The conflicting requirements on frequency deviation for the randomness of TRNG and the reliability of PUF are resolved by an on-chip calibrator, which automatically selects and stores a challenge with a small frequency difference in TRNG mode upon manufacturing and masks unreliable challenges with large frequency difference during PUF enrollment. Leveraging the advantage of custom chip design, the basic delay cell of the reconfigurable ROs is realized by current starved inverter in weak inversion to minimize the power consumption, increase the jitter, and avail its larger process variation. A new lightweight secure mutual authentication protocol is also proposed to effectively thwart machine learning, replay and man-in-the-middle attacks using only the underlying TRNG and PUF without requiring any other security primitives. The proposed TRNG-PUF design is prototyped with a standard 40 nm 1.1 V CMOS process. It occupies a small footprint of 24,$316~\pmb {\mu m^{2}}$. Measured results of the packaged chips show an average energy efficiency of 7.42 pJ/bit in TRNG operation and 0.10 pJ/bit in PUF operation. The bitstreams generated by the test chips passed NIST SP 800-22 and 90B tests, autocorrelation test, and FFT test.
Yuan Cao 0003, Wanyi Liu, Jing Ye 0001, Chip-Hong Chang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Guest Editorial Special Issue on the Asian Hardware Oriented Security and Trust Symposium (AsianHOST 2022)
abstract
Asian Hardware Oriented Security and Trust Symposium (AsianHOST) is an annual symposium that aims to facilitate the rapid growth of hardware-based security research and development. Hardware security is a fashionable research area in both industry and academia. Its scope is consistently growing to embrace secure design, manufacturing, and deployment of modern and emerging interoperable computing, communication, storage devices, and circuits and systems. The 7th Asian Hardware Oriented Security and Trust Symposium (AsianHOST 2022) was held in hybrid mode on December 14–16 in Singapore. Among all the accepted contributions presented at the conference, a subset of top-rated articles was selected and invited for this Special Issue. The invited articles included extended new technical contributions and results and went through a peer-review process consisting of expert reviewers in the related topics. A brief description of the selected articles is as follows.
Chip-Hong Chang, Pingqiang Zhou, Yuan Cao 0003, Qiang Liu 0011
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 A Design of High-Efficiency Coherent Sampling Based TRNG With On-Chip Entropy Assurance
abstract
True Random Number Generator (TRNG) is indispensable in cryptographic algorithms and protocols, and the quality of randomness directly influences the security of cryptographic applications. Multiple theoretical or offline entropy estimation methods have been proposed to evaluate the security of TRNGs, while their ideal assumptions commonly cannot be satisfied due to the perturbation of operating conditions at runtime, which makes it difficult to achieve sufficient entropy for the output of TRNGs in practice. Moreover, the output bitrate of TRNG is another fundamental concern during TRNG practical applications, while popular elementary oscillator-based structure commonly has relatively low output bitrate due to the inherent low sensitivity of entropy extraction to jitter (source of randomness). In this paper, we aim to design a TRNG satisfying both practical security (i.e., on-chip entropy assurance) and high output bitrate simultaneously. In particular, an improved stochastic model and a measurement method are established to quantify the entropy of coherent sampling based TRNG. Moreover, an on-chip entropy assurance module is provided to realize the robustness of the proposed design under various operating conditions. We implement the proposed TRNG in a simulation platform and ASIC chips (with SMIC 130 nm CMOS technology). Experimental results indicate that the generated data has sufficient entropy ($\geq 0.999$per bit) under various operating conditions. In addition, all the output can pass the NIST SP800-22 and AIS 31 statistical tests with an output bitrate of 4.2 Mbps, which is equivalent to 2 orders of magnitude faster than that of the elementary oscillator-based TRNG.
Tianyu Chen 0016, Shijie Jia 0001, Yuan Cao 0003, Wei Wang 0314, Jing Yang 0032, Jingqiang Lin 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 A Voltage Template Attack on the Modular Polynomial Subtraction in Kyber
abstract
Kyber is one of the four final Key Encapsulation Mechanism (KEM) competitors of the National Institute of Standards and Technology PostQuantum Cryptography standardization competition. This paper reveals the vulnerability of Kyber under a voltage template side channel attack: the modular polynomial subtraction operation in Kyber.CCAKEM.Dec. In this paper, by splicing data under different selected ciphertexts, a small number of traces are required to recover the secret key. Experiments show that the recovering accuracy of secret key achieves 100% when using 330 traces, and it still achieves 98% when only using 44 traces.
Jianan Mu, Zongyue Wang, Jing Ye 0001, Junfeng Fan, Huawei Li 0001, Xiaowei Li 0001, Yuan Cao 0003
ASP-DAC9
2022 Nearest Neighbor Classifier with Margin Penalty for Active Learning
Yuan Cao 0003, Zhiqiao Gao, Jie Hu 0025, Jinpeng Chen 0001
ICONIP (1)1
2022 Sequential Intention-aware Recommender based on User Interaction Graph
abstract
The next-item recommendation problem has received more and more attention from researchers in recent years. Ignoring the implicit item semantic information, existing algorithms focus more on the user-item binary relationship and suffer from high data sparsity. Inspired by the fact that user's decision-making process is often influenced by both intention and preference, this paper presents a SequentiAl inTentiOn-aware Recommender based on a user Interaction graph (Satori). In Satori, we first use a novel user interaction graph to construct relationships between users, items, and categories. Then, we leverage a graph attention network to extract auxiliary features on the graph and generate the three embeddings. Next, we adopt self-attention mechanism to model user intention and preference respectively which are later combined to form a hybrid user representation. Finally, the hybrid user representation and previously obtained item representation are both sent to the prediction modul to calculate the predicted item score. Testing on real-world datasets, the results prove that our approach outperforms state-of-the-art methods.
Jinpeng Chen 0001, Yuan Cao 0003, Kaimin Wei
ICMR2
2022 Area, Time and Energy Efficient Multicore Hardware Accelerators for Extended Merkle Signature Scheme
abstract
This paper addresses a barrier that prevents the timely adoption of post-quantum signature algorithms, such as the eXtended Merkle Signature Scheme (XMSS), due to its lack of fast, cost-effective and energy-efficient hardware accelerators. Two new architectures that use more than one hash core are proposed for the first time to significantly reduce the latency of two bottleneck XMSS operations, namely key generation and signature generation, for which the speed of existing hardware accelerators is still apparently inadequate. The first proposed multi-core design uses block RAM and a simplified data flow to maximize the use of$p$hash cores concurrently in three major sequential stages of computation, i. e., Winternitz One-time Signature (WOTS), L-tree and Merkle tree. The second proposed multi-core design adds a dedicated hash core for tree hashing in the L-tree and Merkle tree while keeping the$p$hash cores solely for chain hashing in WOTS. The dedicated hash core leapfrogs between the L-tree and Merkle tree and computes concurrently with the$p$hash cores to keep the$p+1$hash cores active most of the time while minimizing the storage requirement and energy consumption. Both designs are implemented on a 28 nm ATRIX-7 FPGA chip. Experimental results show that both proposed accelerators with$p=8$operate at a much faster speed and consume significantly less hardware resources and energy than all existing XMSS accelerators. Specifically, they are$\sim 8\times $and$\sim 6\times $faster than the fastest reported design in key generation and signature generation operations, respectively.
Yuan Cao 0003, Yanze Wu, Lan Qin, Chip-Hong Chang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 An Efficient Full Hardware Implementation of Extended Merkle Signature Scheme
abstract
This paper presents a full hardware implementation of the eXtended Merkle Signature Scheme (XMSS), a NIST approved and IETF RFC specified post-quantum cryptography (PQC) algorithm. An optimized node traversal is proposed to enable efficient memory utilization without compromising the computational latency of the L-tree and Merkle tree construction, which are two key components used for the compression of the Winternitz One-Time Signature (WOTS) public key in XMSS. The computation of the authentication path during signature generation has also been significantly sped up by our proposed hardware implementation of the Buchmann, Dahmen, and Schneider (BDS) algorithm. Our implementation has completely avoided the use of block random-access memory, which is known to be vulnerable to side-channel attacks. The memory requirement has been highly optimized for implementation with small flip-flop chains and register counters as pointers for fast data access. To the best of our knowledge, this is the first full hardware implementation of all threekey generation,signingandverificationoperations of XMSS. The design has been prototyped and evaluated on a 28 nm FPGA platform to demonstrate its performance improvements over the most efficient software and hardware/software co-design methods reported to date. Specifically, it increases the computational efficiency of the best reported XMSS implementation forkey generationandsignature generationby about 20% and 50%, respectively. It can also run at 10% higher clock speed than the fastest hardware implementation ofsignature verificationin FPGA with 8% lower hardware resource utilization.
Yuan Cao 0003, Yanze Wu, Wen Wang 0007, Jing Ye 0001, Chip-Hong Chang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 A New Energy-Efficient and High Throughput Two-Phase Multi-Bit per Cycle Ring Oscillator-Based True Random Number Generator
abstract
Oscillator-based elementary true random number generator (TRNG) uses a slow jittery ring oscillator (RO) to sample a fast RO. The ROs are always on but most of the oscillatory cycles of the fast RO are not sampled into random bits. In this paper, a new lightweight TRNG design is proposed to minimize the power wasted by the superfluous oscillations. Random bits are extracted from both phases of the slow ROs to increase the throughput and the fast RO is activated only during the narrow transition time difference between two symmetrically designed slow ROs. The slow jittery ROs are implemented using current starved inverters biased in the weak inversion region to reduce their power consumption. Their jitter amplitudes are increased by lowering the oscillation frequency and reducing the drain current of the transistors. The narrow jittery pulse generated by the differential pair of slow ROs is quantized by the fastest three-stage RO. Two random bits from each phase of the jittery ROs can be extracted by using a gigahertz dynamic toggled D flip-flop counter to count the number of oscillatory cycles of the fast RO. The proposed TRNG is fabricated in a standard 65 nm 1.2 V CMOS process. Measurement results of the fabricated chips show that the proposed TRNG consumes merely$260~\mu \text{W}$at a bit rate of 52 Mbps. It outperforms the state-of-art on-chip jitter-based TRNGs with the best figure-of-merit of 5 pJ/bit and the smallest footprint of$366~\mu \text{m}^{2}$. Its generated bit sequence passes the statistical randomness tests including National Institute of Standards and Technology (NIST) test, Auto Correlation Factor (ACF) test and bias. The mean redundancy of the ten tested chips is measured to be less than 10−5bit/symbol.
Yuan Cao 0003, Xiaojin Zhao, Wenhan Zheng, Chip-Hong Chang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 A Lightweight Full Entropy TRNG With On-Chip Entropy Assurance
abstract
True random number generator (TRNG) as one essential hardware primitive is widely used in cryptography, Monte Carlo simulation, and gambling. To evaluate the security of TRNG, the entropy of the TRNG’s output is usually estimated by the stochastic model in theory or measured off-chip after fabrication. However, the sufficiency of entropy is difficult to be guaranteed in practice due to the facts: 1) the inaccuracy of the model-based jitter measurement method; 2) the variations of the chip manufacturing process and operating environments (such as supply voltage and temperature); and 3) malicious attacks. In this work, we design a novel TRNG architecture with on-chip entropy assurance to properly solve practical security problems. In the design, we propose an on-chip entropy estimator for measuring independent jitter to quantify true randomness, which enables continuous monitoring of TRNG at runtime. Furthermore, with the cooperation of the proposed on-chip entropy estimator and a rational self-adaptive mechanism, the designed TRNG can steadily generate bitstreams with sufficient entropy (≥ 0.999 per bit) against PVT variations. We implement the TRNG architecture in FPGAs with different technology nodes (45 and 65 nm) and SMIC 130 nm chips. Experimental results validate that the designed TRNG has an excellent performance in terms of technology independence and environmental robustness. The generated bitstreams pass the NIST SP800-22 and Diehard statistical test suites successfully without any post-processing.
Tianyu Chen 0016, Jingqiang Lin 0001, Yuan Cao 0003, Jiwu Jing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 An All-MOSFET Voltage Reference-Based PUF Featuring Low BER Sensitivity to VT Variations and 163 fJ/Bit in 180-nm CMOS
abstract
In this article, a novel subthreshold voltage reference (VR)-based physical unclonable function (PUF) is presented. With two native nMOS transistors stacked on top for providing bias current and two bottom low threshold voltage (LVT) nMOS transistors forming the self-cascode MOSFETs structure, a 4T all-MOSFET VR is proposed, which features low power consumption and high stability under largely varied VT conditions for the wide range of Internet of Things (IoT) applications. By integrating a pair of the proposed VRs and a digital voltage comparator in each PUF cell, the mismatched output voltages of the VR pair can be locally compared and digitized immediately with the system's power on, leading to ultrashort signal path and maximized immunity to the influence of temporal noise. Fabricated using standard 0.18- μm CMOS process, the proposed PUF design is validated based on extensive measurement results of 20 PUF chips. By passing the widely exploited bias test, National Institute of Standards and Technology (NIST) test and autocorrelation function (ACF) test, the proposed PUF's excellent randomness is well-verified. In addition, the uniqueness is measured to be 49.92%, and the bit error rate (BER) sensitivities in terms of BER per 10 °C and BER per 0.1 V are averaged and reported to be 0.39% and 0.26%, for the temperature range of -40 °C-120 °C and supply voltage range of 1.2-1.8 V, respectively. Moreover, by operating the proposed implementation at a throughput of 50 Mb/s, the measured overall energy consumption is reported to be as low as 163 fJ/bit.
Peizhou Gan, Xiaojin Zhao, Yuan Cao 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Survey: Hardware Trojan Detection for Netlist
abstract
The development of integrated circuit technology is accompanied by potential threats. Malicious modifications to circuits, known as hardware Trojans, are major security concerns. This paper gives a survey of hardware Trojan detection methods towards gate-level netlists. The detection methods are divided into search-based, threshold-based, and machine learning-based ones. This paper compares and analyzes existing works from aspects of feature selection, data balancing techniques, classification criterion, detection range. The experimental results are also selected for comparison.
Yipei Yang, Jing Ye 0001, Yuan Cao 0003, Jiliang Zhang 0002, Xiaowei Li 0001, Huawei Li 0001, Yu Hu 0001
ATS3
2020 Optimization Space Exploration of Hardware Design for CRYSTALS-KYBER
abstract
Public key cryptography is important in the global communication digital infrastructure. However, the emergence of quantum computer and Shor algorithm has greatly threatened the security of public key cryptography. The CRYSTALS-KYBER, as a lattice-based KEM algorithm, passed three rounds of a global solicitation for post-quantum cryptography algorithms held by the National Institute of Standards and Technology (NIST). This paper explores the implementation and optimization space of hardware design according to CRYSTALS-KYBER algorithm. We analyze its software code and try different strategies to optimize the hardware implementation, and conduct comparative analysis in terms of area and speed. The experimental results show that the performance can be greatly improved by moderately optimizing the loops. In comparison with optimal results of the work [12], our optimizations improve the performance by up to 74.6% for encapsulation algorithm and 54.4% for decapsulation algorithm.
Zhiteng Chao, Jing Ye 0001, Wen Wang 0007, Yuan Cao 0003, Xiaowei Li 0001, Huawei Li 0001
ATS5
2020 A 30fJ/b Current-Biased Inverter Based RO TRNG with High Temperature and Supply Voltage Stabilities
abstract
In this paper, we present an ultra-low power true random number generator (TRNG) based on ring oscillator (RO) with current-biased inverters. Random numbers are extracted by symmetrically designed arbiter depending on the jitter-noise-caused phase difference of a pair of ROs. Using a bias transistor shared by 16 inverters (8 for each RO), we can obtain a virtual power supply lower than VDD(denoted as V VDD) and a subthreshold bias current to increase the oscillation frequency while ensuring high randomness. In addition, power consumption is significantly lowered by using V VDDto reduce the amplitude of oscillation. Moreover, V VDDcan be designed with high temperature and supply voltage stability by optimizing the transistor sizes, which allows the proposed implementation to generate random numbers and pass both NIST and auto-correlation test suites over wide ranges of supply voltage (1.0V~1.4V) and temperature (0°C~120°C). The proposed TRNG is implemented using standard 65nm CMOS process, with the bit generation rate and energy efficiency reported to be 169Mbps and 30fJ/bit, respectively (1.2V and 27°C).
Shengquan Liang, Wenhan Zheng, Yuan Cao 0003, Xiaojin Zhao
ISCAS3
2020 A PUF-Based Data-Device Hash for Tampered Image Detection and Source Camera Identification
abstract
With the increasing prevalent of digital devices and their abuse for digital content creation, forgeries of digital images and video footage are more rampant than ever. Digital forensics is challenged into seeking advanced technologies for forgery content detection and acquisition device identification. Unfortunately, existing solutions that address image tampering problems fail to identify the device that produces the images or footage while techniques that can identify the camera is incapable of locating the tampered content of its captured images. In this paper, a new perceptual data-device hash is proposed to locate maliciously tampered image regions and identify the source camera of the received image data as a non-repudiable attestation in digital forensics. The presented image may have been either tampered or gone through benign content preserving geometric transforms or image processing operations. The proposed image hash is generated by projecting the invariant image features into a physical unclonable function (PUF)-defined Bernoulli random space. The tamper-resistant random PUF response is unique for each camera and can only be generated upon triggered by a challenge, which is provided by the image acquisition timestamp. The proposed hash is evaluated on the modified CASIA database and CMOS image sensor-based PUF simulated using 180 nm TSMC technology. It achieves a high tamper detection rate of 95.42% with the regions of tampered content successfully located, a good authentication performance of above 98.5% against standard content-preserving manipulations, and 96.25% and 90.42%, respectively, for the more challenging geometric transformations of rotation (0 ~ 360°) and scaling (scale factor in each dimension: 0.5). It is demonstrated to be able to identify the source camera with 100% accuracy and is secure against attacks on PUF.
Yuan Cao 0003, Chip-Hong Chang
IEEE Trans. Inf. Forensics Secur.2
2020 Ed-PUF: Event-Driven Physical Unclonable Function for Camera Authentication in Reactive Monitoring System
abstract
As surveillance footage plays an increasingly significant role in law enforcement, it is imperative to ensure the integrity of recorded video data and the authenticity of its originator, and instill situation awareness into these monitoring systems with a fidelity record of the incidents. Unfortunately, existing frame-based networked surveillance systems could only partially fulfill these requirements. The emerging Dynamic Vision Sensor (DVS) sheds new light on solving this problem with its completely different sensor design, i.e., DVS responds only to temporal intensity change and records only sparse asynchronous address-events with precise timing information. Motivated by the reduced data size of activities and the prevention of privacy intrusion of subjects under surveillance as well as other appealing attributes, this work introduces the first event-driven physical unclonable function (Ed-PUF) system to fill the forensic gap of simultaneously authenticating the event data integrity and source camera identity for reactive monitoring by DVS camera. New DVS sensor architecture is proposed with negligible modifications made to the original DVS pixel. The Ed-PUF response bit can only be triggered by and uniquely dependent on the asynchronous addressed event without being interfered by the simultaneous firing of other address events. Address event streams are securely transmitted with an event package tag created by a keyed hash-based message authentication code with the key being the Ed-PUF response. A secure protocol to authenticate the identity of DVS camera and the integrity of address events transmitted through cellular network is also proposed. A camera lock is embedded to protect against severing and splicing the inter-chip connectivity within the camera for raw PUF responses. The proposed system is evaluated using raw PUF data obtained by post-layout Monte Carlo simulation in UMC 180nm technology and real event stream captured by a DVS camera. The proposed Ed-PUF has been demonstrated to have excellent uniqueness, randomness and reliability. Collision test is also conducted to show that the quality of DVS imaging is not compromised. Besides keeping the hardware/power/timing overheads low, the proposed scheme is also analyzed to be resilient against multiple attack scenarios.
Xiaojin Zhao, Takashi Sato 0001, Yuan Cao 0003, Chip-Hong Chang
IEEE Trans. Inf. Forensics Secur.4
2020 A 1036-F2/Bit High Reliability Temperature Compensated Cross-Coupled Comparator-Based PUF
abstract
In this article, a compact physical unclonable function (PUF) based on cross-coupled comparator is presented. Featuring a positive feedback response generation mechanism, the mismatch in analog signals between the cross-coupled transistor pair is quickly amplified to prevent its polarity from flipping by the temporal noise. The rapid enlargement of noise margin by the sense amplifier also contributes to stabilizing the response against supply voltage variations. To improve its temperature stability, the counteracting effect of complementary-to-absolutetemperature (CTAT) and proportional-to-absolute-temperature (PTAT) drives are considered in sizing the bit cell transistors. The proposed design is fabricated in a standard 65-nm CMOS process. The bit cell occupies an area of only 4.38 μm2(i.e., 1036 F2), and the overall PUF chip consumes 2.98 pJ/bit at the throughput of 8 Mb/s, of which only 1.61 pJ/bit is due to the PUF's core. With the uniqueness measured to be 49.53%, the unpredictability of the fabricated PUF chips is validated by autocorrelation function and NIST randomness tests. Compared with the state-of-the-art implementations, the proposed PUF has the lowest native response instability of 1.46% with 500 repeated PUF readouts at 27 °C and 1.2 V. By varying the operating temperature from -50 °C to 150 °C in a step size of 10 °C and the supply voltage from 1.0 to 1.4 V in a step size of 0.1 V simultaneously, the average reliability of the proposed PUF obtained from the 2-D plot of all operating conditions is found to be 96.87% without correction and 99.31% with spatial majority voting (SMV).
Yiheng Wu, Xiaojin Zhao, Yuan Cao 0003, Chip-Hong Chang
IEEE Trans. Very Large Scale Integr. Syst.4
2019 High-Speed True Random Number Generator Based on Differential Current Starved Ring Oscillators with Improved Thermal Stability
abstract
True random number generator (TRNG) is a hardware security primitive that has been widely used in cryptography, Monte-Carlo simulation, video games, etc. with increasing significance in security solutions addressing emerging cyber-physical threats. This paper presents a new design of TRNG based on temperature insensitive frequency deviation extracted from a pair of matched current starved ring oscillators (CS-ROs). The proposed TRNG can maintain the entropy rate of the output bit stream over a broad range of temperatures. The CS-RO generates larger jitter noise than the regular ring oscillator, but consumes much less power for the same number of inverter stages. Monte Carlo simulated results based on a commercial 65 nm 1.2V CMOS process technology shows that it can output a random bit sequence at a rate of 485 Mbps with only 1.52 pJ of energy dissipation per bit. The proposed TRNG passes the NIST cryptographic randomness tests and autocorrelation test for the bit streams generated at temperature ranges from −40°C to 120°C.
Yuan Cao 0003, Chip-Hong Chang, Xiaoli Ji
ISCAS2
2019 A Reliable Physical Unclonable Function Based on Differential Charging Capacitors
abstract
Physical Unclonable Function (PUF) is an emerging security primitive for cryptography applications. However, achieving a very high reliability against the environmental variations remains a main challenge in PUF design and a key barrier for its commercialization. This paper presents a new PUF design based on the charging of a symmetric MOS capacitor pair by constant current with cross-coupled positive feedback inverters. The proposed weak PUF features high raw response reliability against variations in power supply and temperature without power-up reset noise and other issues due to the power-down and up of an array of cells. Extensive Monte-Carlo simulations have been performed using a standard 110nm CMOS process technology. The simulated results show an almost ideal uniqueness of 50.03% and superior reliability of 97.70% over a temperature range from 0 °C to 80 °C, and 96.20% with the supply voltage varies from 1.2 V to 1.8 V. The response bit can be generated at a rate of 27.78 Mbps with an average power consumption of 20.86 μW at 1.5V, and the energy consumption is only 750 fJ/bit.
Wei Guo 0018, Chip-Hong Chang, Yuan Cao 0003, Shaojun Wei, Shouyi Yin, Chenchen Deng, Leibo Liu, Fan Zhang 0044
ISCAS4
2019 A Highly Reliable Physical Unclonable Function Based on 2T Voltage Reference and Diode-Clamped Comparator
abstract
In this paper, we present a novel physical unclonable function (PUF) structure based on 2T voltage reference and diode-clamped comparator. The proposed PUF implementation includes an array of the aforesaid 2T voltage references with same transistor size and layout design. Meanwhile, the output voltage variation mainly caused by the CMOS process variation can be well-extracted by the adopted diode-clamped comparator to generate a digital bit stream (256 bit) with excellent randomness. By utilizing the sub-threshold 2T voltage reference's superior stability to the environmental variations (e.g. operating temperature, supply voltage), the proposed implementation exhibits ultra-high reliability against wide-range operating temperature and supply voltage. This design is validated by our extensive post-layout simulations based on 65nm standard CMOS process, the average reliability is reported to be 99.88% and 99.37% for the operating temperature ranging from -40°C to 120°C and the supply voltage ranging from 0.8V to 1.8V, respectively. In addition, the overall power consumption is reduced to 3.1μW at a throughput of 10 Mb/s. Moreover, the superior unpredictability (i.e. randomness) of this design is verified by passing both the auto-correlation test and NIST randomness test.
Xiaojin Zhao, Yuan Cao 0003
ISCAS3
2019 Umbrella: Enabling ISPs to Offer Readily Deployable and Privacy-Preserving DDoS Prevention Services
abstract
Defending against distributed denial of service (DDoS) attacks on the Internet is a fundamental problem. However, recent industrial interviews with over 100 security experts from more than ten industry segments indicate that DDoS problems have not been fully addressed. The reasons are twofold. On one hand, many academic proposals that are provably secure witness little real-world deployment. On the other hand, the operation model for existing DDoS-prevention service providers (e.g., Cloudflare, Akamai) is privacy invasive for large organizations (e.g., government). In this paper, we present Umbrella, a new DDoS defense mechanism enabling Internet service providers to offer readily deployable and privacy-preserving DDoS prevention services to their customers. At its core, Umbrella develops a multi-layered defense architecture to defend against a wide spectrum of DDoS attacks. In particular, the flood throttling layer stops amplification-based DDoS attacks; the congestion resolving layer, aiming to prevent sophisticated attacks that cannot be easily filtered, enforces congestion accountability to ensure that legitimate flows are guaranteed to receive their fair shares regardless of attackers' strategies; and finally the user-specific layer allows DDoS victims to enforce self-desired traffic control policies that best satisfy their business requirements. Based on Linux implementation, we demonstrate that Umbrella is capable to deal with large-scale attacks involving millions of attack flows, meanwhile imposing negligible packet processing overhead. Further, our physical test bed experiments and large-scale simulations prove that Umbrella is effective to mitigate various DDoS attacks.
Zhuotao Liu, Yuan Cao 0003, Min Zhu 0001
IEEE Trans. Inf. Forensics Secur.2
2019 Managing Recurrent Virtual Network Updates in Multi-Tenant Datacenters: A System Perspective
abstract
With the advent of software-defined networking, network configuration through programmable interfaces becomes practical, leading to various on-demand opportunities for network routing update in multi-tenant datacenters, where tenants have diverse requirements on network routings such as short latency, low path inflation, large bandwidth, high reliability, etc. Conventional solutions that rely on topology search coupled with an objective function to find desired routings have at least two shortcomings: ${\sf (i)}$(i) they run into scalability issues when handling consistent and frequent routing updates and ${\sf (ii)}$(ii) they restrict the flexibility and capability to satisfy various routing requirements. To address these issues, this paper proposes a novel search and optimization decoupled design, which not only saves considerable topology search costs via search result reuse, but also avoids possible sub-optimality in greedy routing search algorithms by making decisions based on the global view of all possible routings. We implement a prototype of our proposed system, OpReduce, and perform extensive evaluations to validate its design goals.
Zhuotao Liu, Yuan Cao 0003, Xuewu Zhang 0001, Changping Zhu, Fan Zhang 0010
IEEE Trans. Parallel Distributed Syst.2
2018 A Fully Digital Physical Unclonable Function Based Temperature Sensor for Secure Remote Sensing
abstract
Turnkey solutions that combine energy-efficient remote sensing and secure communication of telemetry are desirable in data collection, risk control and situation appraisal with the large scale deployment of resource constrained Internet of Things devices. In this paper, a new low-cost physical unclonable function (PUF) based temperature sensor for secure remote temperature sensing is proposed. The design exploits the approximately linear positive temperature coefficient of CMOS inverter in super-threshold operation to calibrate the running frequency of ring oscillator (RO) in a reconfigurable RO PUF at different temperature. The RO frequency corresponding to the sensed temperature is fed into a randomizer seeded by the input challenge to select new RO pairs for comparison to generate a random, unique and physically unclonable digital tag, which is valid for a selected input challenge to a target device at a particular temperature. Using only standard logic cells and a very simple structure, the proposed temperature sensor can be easily implemented on FPGA and integrated into other digital systems. It protects the integrity of the sensed information by preventing falsified sensor data and masquerade sensing node. The FPGA implementation of our proposed design has demonstrated the feasibility of making a trust temperature telemetry system out of PUF.
Yuan Cao 0003, Yunyi Guo, Benyu Liu, Min Zhu 0001, Chip-Hong Chang
ICCCN1
2018 A Sub-pico Joules Per Bit Robust Physical Unclonable Function Based on Subthreshold Voltage References
abstract
Low power, lightweight and robust physical unclonable function (PUF) is a sought-after for IoT device identification/authentication. This paper presents a low power PUF design with high reliability against temperature and supply voltage variations. A response bit is extracted by comparing a pair of identically designed subthreshold voltage references. The voltage difference due to device mismatch is digitized and registered in a bidirectional counter, which can be used to identify and filter out the unstable response bits. The readout circuit works in tandem with the proposed double sampling technique to reduce the bias of components that are not the main entropy source of response bits. The proposed design is evaluated by extensive simulation using standard 65 nm CMOS process. It consumes merely 0.16 pJ/bit. The simulated uniqueness is an almost ideal 50.03%. Due to the intrinsic stability of voltage references, the reliability of its native response is 98.17% for the supply voltage variation from 1 V to 1.4 V and 97.60% for the temperature variation from 0 ° C to 80 °C. The generated response bitstream has passed both the autocorrelation test and NIST randomness test.
Yuan Cao 0003, Chip-Hong Chang, Wenhan Zheng, Xiaojin Zhao
ISCAS1
2017 A novel smoothness-based interpolation algorithm for division of focal plane Polarimeters
abstract
In this paper, we present a novel smoothness-based interpolation algorithm for the division of focal plane Polarimeters (DoFP). By calculating the divided blocks' variance that represents their local smoothness, the proposed algorithm well-balances between the traditional bilinear and bicubic interpolation algorithms. In addition, compared with the previously reported gradient-based interpolation algorithm which only indicates the image's directional change along 0°, 90°, 45° and 135°, the presented smoothness-based algorithm covers the variations along all the possible directions, leading to more accurate selection between the bilinear and bicubic algorithms. According to our extensive simulation results, the proposed implementation exhibits the lowest mean square error (MSE) for the test images among all the previously reported algorithms, including bilinear, bicubic and gradient-based interpolation algorithms.
Jieyun Zhang, Wen Bin Ye 0001, Ashfaq Ahmed, Zhurui Qiu, Yuan Cao 0003, Xiaojin Zhao
ISCAS5
2016 An energy-efficient subthreshold level shifter with a wide input voltage range
abstract
The level shifters are crucial primitives in the multi-supply voltage circuits and systems. In this paper, an energy-efficient level shifter is proposed to achieve the conversion from the subthreshold voltage to the above threshold voltage. It is a hybrid structure consisting of the Wilson current mirror and the cross-coupled level shifter. By addressing the voltage drop issue of the level shifter based on the Wilson current mirror, the leakage power is significantly reduced, with the advantage of wide input voltage range for the Wilson current mirror level shifter well-preserved. In addition, the multi-threshold CMOS (MTCMOS) technology is employed to provide more flexibility for our ultra-low power design. The reported simulation results using 65 nm CMOS process validate our proposed implementation and an ultra-low power consumption of 19.44 fJ per conversion from 0.2 V to 1.2 V at 1 MHz is achieved without the need of any intermediate power supply.
Yuan Cao 0003, Wen Bin Ye 0001, Xiaojin Zhao, Peigang Deng
ISCAS1
2016 A compact ultra-low power physical unclonable function based on time-domain current difference measurement
abstract
In this paper, we present a novel physical unclonable function (PUF) based on time-domain current difference measurement. By employing the aforesaid simplified current-mode PUF architecture, the proposed implementation completely removes the need of complex error correction circuitry, which is widely adopted in the previously demonstrated implementations. This leads to significant reduction of both the overall power consumption and the required chip area. In addition, the proposed implementation exhibits a superior bit error rate (BER) as low as 0 for the typical case scenario and 1.56% for the worst case scenario, respectively. Featuring an ultra-low power consumption of 11.29μW and an averaged silicon area of 13310μm2, the proposed implementation is validated by our reported extensive post-layout simulation results with UMC 0.18μm standard complementary-metal-oxide-semiconductor (CMOS) technology.
Shibang Lin, Yuan Cao 0003, Xiaojin Zhao, Xiaofang Pan
ISCAS2
2015 A Low-Power Hybrid RO PUF With Improved Thermal Stability for Lightweight Applications
abstract
Ring oscillator (RO)-based physical unclonable function (PUF) is resilient against noise impacts, but its response is susceptible to temperature variations. This paper presents a low-power and small footprint hybrid RO PUF with a very high temperature stability, which makes it an ideal candidate for lightweight applications. The negative temperature coefficient of the low-power subthreshold operation of current starved inverters is exploited to mitigate the variations of differential RO frequencies with temperature. The new architecture uses conspicuously simplified circuitries to generate and compare a large number of pairs of RO frequencies. The proposed nine-stage hybrid RO PUF was fabricated using global foundry 65-nm CMOS technology. The PUF occupies only 250 μm2of chip area and consumes only 32.3 μW per challenge response pair at 1.2 V and 230 MHz. The measured average and worst-case reliability of its responses are 99.84% and 97.28%, respectively, over a wide range of temperature from -40 to 120 °C.
Yuan Cao 0003, Le Zhang 0001, Chip-Hong Chang, Shoushun Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 A Cluster-Based Distributed Active Current Sensing Circuit for Hardware Trojan Detection
abstract
The globalization of integrated circuits (ICs) design and fabrication has given rise to severe concerns on the devastating impact of subverted chip supply. Hardware Trojan (HT) is among the most dangerous threats to defend. The dormant circuit inserted stealthily into the chip by the advisory could steal the confidential information or paralyze the system connected to the subverted chip upon the HT activation. This paper presents a transient power supply current sensor to facilitate the screening of an IC for HT infection. Based on the power gating scheme, it converts the current activity on local power grid into a timing pulse from which the timing and power-related side channel signals can be externally monitored by the existing scan test architecture. Its current comparator threshold can be calibrated against the quiescent current noise floor to reduce the impacts of process variations. Postlayout statistical simulations of process variations are performed on the ISCAS'85 benchmark circuits to demonstrate the effectiveness of the proposed technique for the detection of delay-invariant and rarely switched HTs. Compared with the detection error rate of a 4-bit counter-based HT reported by an existing HT detection method using the path delay fingerprint, our method shows an order of magnitude improvement in the detection accuracy.
Yuan Cao 0003, Chip-Hong Chang, Shoushun Chen
IEEE Trans. Inf. Forensics Secur.1
2013 Cluster-based distributed active current timer for hardware Trojan detection
abstract
With the globalization of integrated circuit (IC) design and fabrication, there is a growing concern on the devastating impact of subverted chip supply. This paper presents a current sensing circuit that converts the current activity on local power grid to a timing pulse to detect if an IC is Trojan-infected. This new approach increases the Trojan detection sensitivity by combining the switching activity and path sensitization abnormalities into a single side-channel signal that can be easily monitored by existing scan test structure. One main advantage of the proposed regional Trojan detector is that the current comparator threshold can be calibrated against the quiescent current noise floor to reduce the impacts of process variations. Experiments are performed on a Trojan-infected benchmark circuit to demonstrate the feasibility of the proposed technique.
Yuan Cao 0003, Chip-Hong Chang, Shoushun Chen
ISCAS1