VLDB 2026 Research / reviewers in the wild / expert
Chip-Hong Chang
dblp:95/137 · also Chip Hong Chang
· DBLP profile ↗
153ranked-venue papers
17as first author
42since 2021 · last 2026
0000-0002-8897-6176ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 109 · 8 first-author · 24 since 2021Security and privacy · 24 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Computer networks · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QT-PUF: Quantum Tunneling Leakage Based PUF for Implantable IoMT Devices
Yueqi Ma, Vivek Mohan, Chip-Hong Chang, Emmanuel M. Drakakis |
ISCAS | 3 |
| 2026 | SLAPE: Secure Lightweight Authentication for Privacy-Enhanced Federated Learning in Industrial Internet of ThingsabstractFederated learning (FL) has emerged as a promising technique in the Industrial Internet of Things (IIoT) by enabling distributed devices to collaboratively train models without sharing raw data. In FL, ensuring data privacy and secure authentication becomes essential due to the sensitivity of industrial data and the potential for adversarial attacks. This paper highlights a security flaw in a recently proposed FL authentication protocol designed for IIoT environments. Specifically, the scheme is analyzed to be susceptible to public-key replacement attacks. We propose a secure, lightweight authentication scheme for privacy-enhanced federated learning (SLAPE) to address vulnerabilities in participant registration, group key distribution, local data training, and aggregation processes. SLAPE leverages the Elliptic Curve Cryptography with the Chinese Remainder Theorem to support malicious group member traceability, revocation of compromised identities, and efficient batch verification of multiple messages. It effectively resists Type-I attacks that previous schemes could not, while also incorporating forward and backward security essential for IIoT applications. We rigorously demonstrate SLAPE’s resilience against prevalent threats through both formal and informal analyses. Our evaluation results indicate that SLAPE demonstrably enhances the security and privacy of existing schemes, with improvements in computational efficiency for both proof generation and verification, while keeping communication overhead relatively low. Shanyao Ren, Jianwei Liu 0001, Chip-Hong Chang, Hanzhou Wang, Dongyu Li |
IEEE Internet Things J. | 3 |
| 2026 | SMARC: A State-Repairing Multi-Agent Resilient Consensus SchemeabstractIn this paper, we analyze a recent algorithm for resilient consensus control in distributed multi-agent systems. While effective in theory, its reliance on security in communication and strong connectivity assumptions limits its practicality in dynamic electronic and cyber-physical systems, such as embedded device networks and uncrewed platforms. To address these limitations, we propose a State-repairing Multi-agent Resilient Consensus (SMARC) scheme to eliminate the need for normal agents to collect trustworthy state values from a fixed bounded threshold of neighbors to achieve reliable consensus. The core innovation of SMARC is a decentralized state-repair mechanism, which enables agents to obtain additional information from their reachable sets to repair unavailable or corrupted state values of malicious or faulty agents, and autonomously adjust their convergence speed. Additionally, without negatively impacting the consensus performance, SMARC employs a noise-masked surface state to protect the initial states of agents from eavesdropping. This approach avoids extra storage and reduces computations without adhering to strict security prerequisites, making it more suitable for resource-constrained electronic systems compared to existing methods. Theoretical convergence and security proofs demonstrate that SMARC can successfully resist passive attacks while ensuring accurate convergence on multi-dimensional data. Most importantly, from small- to large-scale networks, SMARC achieves a three- to four-fold increase in convergence speed compared to the most competitive recent state-of-the-art resilient consensus algorithm in both passive and active attacks. A prototype electronic of a MAS was also built using six Raspberry Pi devices to validate its performance and robustness in practical environments. Shanyao Ren, Chip-Hong Chang, Jianwei Liu 0001, Dongyu Li |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | RAFS: Reversible Identity-Anonymization Face Swapping for Provenance TrackingabstractFace-swapping technologies have rapidly emerged as a mainstream AI service across entertainment, social media, and virtual platforms. While current face-swapping methods offer highly realistic results, they also introduce significant privacy risks, as most current approaches require clear target faces, exposing users’ identities. Moreover, the absence of built-in authorization and forensic mechanisms renders these systems incapable of tracing or verifying manipulated content, raising critical issues over accountability and potential misuse. To address these challenges, we propose a privacy-preserving and forensics-enabled face-swapping framework that simultaneously safeguards user identity and enables robust post-hoc face recovery. Instead of relying on visible target faces, our method operates on non-facial target images, fundamentally preventing identity exposure at the source. To ensure provenance traceability, we embed the target face’s features into non-facial regions of the generated image via an imperceptible and reversible encoding scheme. To further enhance robustness, we introduce a mask-distortion simulation layer that bridges pre-/post-swap mask discrepancies and stabilizes face recovery under perturbations. Extensive experiments demonstrate that our method produces realistic face-swapped images without revealing the original identity, while enabling high-fidelity recovery under various adversarial conditions—validating its effectiveness in both privacy protection and forensic traceability. Jiancheng Li, Peipeng Yu, Chip-Hong Chang, Zhangjie Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | DFREC: DeepFake Identity Recovery Based on Identity-Aware Masked AutoencoderabstractRecent advances in deepfake forensics have primarily focused on improving the classification accuracy and generalization performance. Despite enormous progress in detection accuracy across a wide variety of forgery algorithms, existing algorithms lack intuitive interpretability and identity traceability to help with forensic investigation. In this paper, we introduce a novel DeepFake Identity Recovery scheme (DFREC) to fill this gap. DFREC aims to recover the pair of source and target faces from a deepfake image to facilitate deepfake identity tracing and reduce the risk of deepfake attacks. It comprises three key components: an Identity Segmentation Module (ISM), a Source Identity Reconstruction Module (SIRM), and a Target Identity Reconstruction Module (TIRM). The ISM segments the input face into distinct source and target face information, and the SIRM reconstructs the source face and extracts latent target identity features with the segmented source information. The background context and latent target identity features are synergetically fused by a Masked Autoencoder in the TIRM to reconstruct the target face. We evaluate DFREC on different high-fidelity face-swapping attacks on FaceForensics++, CelebaMegaFS, FFHQ-E4S, and Celeb-DFv2 datasets, which demonstrate its superior recovery performance over state-of-the-art deepfake recovery algorithms. In addition, DFREC is the only scheme that can recover both pristine source and target faces directly from the forgery image with high fidelity. Peipeng Yu, Jianwei Fei, Zhihua Xia, Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | Adversarial Attacks Against Deep Learning-Based Radio Frequency Fingerprint IdentificationabstractRadio frequency fingerprint identification (RFFI) is an emerging technique for the lightweight authentication of wireless Internet of things (IoT) devices. RFFI exploits deep learning models to extract hardware impairments to uniquely identify wireless devices. Recent studies show deep learning-based RFFI is vulnerable to adversarial attacks. However, effective adversarial attacks against different types of RFFI classifiers have not yet been explored. In this paper, we carried out a comprehensive investigations into different adversarial attack methods on RFFI systems using various deep learning models. Three specific algorithms, fast gradient sign method (FGSM), projected gradient descent (PGD), and universal adversarial perturbation (UAP), were analyzed. The attacks were launched to LoRa-RFFI and the experimental results showed the generated perturbations were effective against convolutional neural networks (CNNs), long short-term memory (LSTM) networks, and gated recurrent units (GRU). We further used UAP to launch practical attacks. Special factors were considered for the wireless context, including implementing real-time attacks, the effectiveness of the attacks over a period of time, etc. Our experimental evaluation demonstrated that UAP can successfully launch adversarial attacks against the RFFI, achieving a success rate of 81.7% when the adversary almost has no prior knowledge of the victim RFFI systems. Junqing Zhang, Guanxiong Shen, Alan Marshall 0001, Chip-Hong Chang |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | A Lightweight PUF-Based Secure Group Communication Scheme for Low Altitude Network With Dynamic Group MembershipabstractLow Altitude Network (LAN) has emerged as a critical infrastructure for applications such as surveillance and emergency response. Group communication for LAN offers enhanced energy efficiency and reduced network overhead. However, existing group communication protocols encounter difficulties in managing the rekeying process efficiently for a dynamic group or require computationally expensive public key primitives for shared secret handshaking to overcome this challenge. Moreover, the security of most of the existing protocols relies primarily on the safekeeping of some secrets on the group members' devices. To overcome these limitations, we propose a novel physical unclonable function (PUF)-based lightweight secure group communication protocol. The proposed protocol utilizes a combination of the device's PUF and the one-time pad (OTP) to eliminate secure key storage at both group verifier and prover nodes, and achieves perfect forward secrecy (PFS) by eliminating dependence on static long-term secrets. The proposed protocol also supports efficient group key renewal by using a full binary tree as a secret vault for sharing and updating distributed secrets with the Chinese Remainder Theorem (CRT). Meantime, this data structure also reduces the computation and communication complexity for key renewal to$O(\log _{2}N)$at the cluster head and$O(1)$at the sensor nodes. A comparative analysis shows that the proposed protocol surpasses related protocols in terms of security features and overheads in computation, communication, as well as secret storage requirements. The proposed protocol was also validated by formal security analyses and a physical LAN implementation using Ultra96-V2 boards as cluster nodes. Harishma Boyapally, Wenye Liu, Yongkui Yang, Chip-Hong Chang |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Robust Secure Swap: Responsible Face Swap With Persons of Interest Redaction and Provenance TraceabilityabstractAs AI generative models evolve, face swap technology has become increasingly accessible, raising concerns over potential misuse. Celebrities may be manipulated without consent, and ordinary individuals may fall victim to identity fraud. To address these threats, we propose Secure Swap, a method that protects persons of interest (POI) from face-swapping abuse and embeds a unique, invisible watermark into nonPOI swapped images for traceability. By introducing an ID Passport layer, Secure Swap redacts POI faces and generates watermarked outputs for nonPOI. A detachable watermark encoder and decoder are trained with the model to ensure provenance tracing. Experimental results demonstrate that Secure Swap not only preserves face swap functionality but also effectively prevents unauthorized swaps of POI and detects different embedded model’s watermarks with high accuracy. Specifically, our method achieves a 100% success rate in protecting POI and over 99% watermark extraction accuracy for nonPOI. Besides fidelity and effectiveness, the robustness of protected models against image-level and model-level attacks in both online and offline application scenarios is also experimentally demonstrated. Yunshu Dai, Jianwei Fei, Fangjun Huang, Chip-Hong Chang |
ICML | 4 |
| 2025 | Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake DetectionabstractCurrent Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data, but their potential remains underexplored for deepfake detection due to the misalignment of their knowledge and forensics patterns. To this end, we present a novel framework that unlocks LVLMs’ potential capabilities for deepfake detection. Our framework includes a Knowledge-guided Forgery Detector (KFD), a Forgery Prompt Learner (FPL), and a Large Language Model (LLM). The KFD is used to calculate correlations between image features and pristine/deepfake image description embeddings, enabling forgery classification and localization. The outputs of the KFD are subsequently processed by the Forgery Prompt Learner to construct fine-grained forgery prompt embeddings. These embeddings, along with visual and question prompt embeddings, are fed into the LLM to generate textual detection responses. Extensive experiments on multiple benchmarks, including FF++, CDF2, DFD, DFDCP, DFDC, and DF40, demonstrate that our scheme surpasses state-of-the-art methods in generalization performance, while also supporting multi-turn dialogue capabilities. Peipeng Yu, Jianwei Fei, Xuan Feng 0002, Zhihua Xia, Chip-Hong Chang |
ICML | 6 |
| 2025 | PUF-based Edge DNN Model IP Protection with Self-obfuscation and Publicly Verifiable OwnershipabstractIntellectual Property (IP) violation poses severe threats to Deep Neural Network (DNN) models deployed on easily accessible edge devices for real-time, security- and saff1ety-critical applications. Existing DNN IP protection schemes lack consideration of hardware attack vectors and fall short of the security demands of edge DNNs. In this paper, we leverage hardware-software co-design for pre-emptive protection of DNN models implemented on edge devices against multiple field attacks. A unified credential is generated by projecting a subset of model's weights selected by the physical unclonable function of its hosting device into a random space spanned by its owner ID. The model is trained in such a way that the aff1ine parameters of selected normalization blocks are modulated by this credential to provide accurate inference on valid edge device and owner ID, which allows the ownership to be publicly and black-box verifiable by its users. The deployed inference machine lockdowns automatically upon any attempts to export, modify or retrain the model, thus effectively prevent many forms of infringement, including ownership forgery, model extraction and replication, unauthorized domain adaptation and reimplementation of model on unauthorized devices. We have evaluated these attacks and the ownership proof with two DNN models and four datasets. The experimental results demonstrated that the protected edge DNNs produce high inference fidelity on correct device and valid owner ID, and significantly degraded accuracy on all attacks. Jingdong Jiang, Chip-Hong Chang |
ISCAS | 3 |
| 2025 | Accuracy-preserving Layer Normalization Approximations for Efficient Transformer Hardware AcceleratorsabstractTransformers have emerged as leading models for solving natural language processing (NLP) problems. Layer normalization (LN) has been found to be a throughput and latency bottleneck for Transformer networks. The LN datapath involves sequential processing that is dependent on the input data, and GPU and CPU unfriendly nonlinear square root and reciprocal operations. These make pipelining and hardware simplification challenging without compromising the accuracy of pretrained models. In this paper, we propose a hardware-efficient and accuracy-preserving design for LN approximation. The LN core can be directly plugged into pretrained Transformer networks to accelerate NLP tasks without fine tuning. The FPGA implementation of our design outperforms the state-of-the-art FPGA implementation of LN, with 3.5 × higher throughput per LUT and 1.3 × higher throughput per DSP. Our approximation introduces negligible average and worst-case accuracy drops of only 0.28% and 1.09%, respectively on the main language benchmark tasks. Additionally, a better pipelined design with pairwise variance calculation is proposed to reduce the number of iterations for LN, resulting in a further 27% reduction in latency with a slight increase in hardware resource requirement. Nazim Altar Koca, Anh-Tuan Do, Chip-Hong Chang |
ISCAS | 3 |
| 2025 | Red Bleed: A Pragmatic Near-Infrared Presentation Attack on Facial Biometric Authentication Systems
Chip-Hong Chang |
USENIX Security Symposium | 3 |
| 2025 | HAST: A Hardware-Efficient Spatio-Temporal Correlation Near-Sensor Noise Filter for Dynamic Vision SensorsabstractThe Dynamic Vision Sensor (DVS) is a bio-inspired image sensor which has many advantages such as high dynamic range, high bandwidth, high temporal resolution and low power consumption for Internet of Video Things and Edge Computing applications. However, spuriously generated Background Activity (BA) noise events can significantly degrade the quality of DVS output and cause unnecessary computations throughout the image processing chain, reducing its energy efficiency. Near-sensor filters can mitigate this problem by preventing the BA noise events from reaching downstream stages. In this paper, we propose a novel, hardware-efficient, spatio-temporal correlation filter (HAST) for near-sensor BA noise filtering. It uses compact two-dimensional binary arrays along with simple, arithmetic-free hash-based functions for storage and retrieval operations. This approach eliminates the need to use timestamps for determining the chronological order of events. HAST uses much lower memory and energy compared to other hardware-friendly filters (BAF/STCF) while matching their performance in simulations with standard datasets; for a sensor of resolution$346\times 260$pixels, it requires only 5–18% of their memory, and about 15% of their energy per event for correlation time$\tau $ranging from 1 to 50 ms. The memory and energy gains of the filter increase with sensor resolution. In FPGA implementation, HAST achieves about 29% higher throughput than BAF/STCF while utilizing only about 5% of their memory. The filter parameter values can be chosen by Design Space Exploration (DSE) for optimized performance-resource trade-offs based on application requirements. Pradeep Kumar Gopalakrishnan, Chip-Hong Chang, Arindam Basu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | ASHES '24: Workshop on Attacks and Solutions in Hardware SecurityabstractThe workshop on "Attacks and Solutions in HardwarE Security (ASHES)" welcomes any theoretical and practical works on hardware security, including attacks, solutions, countermeasures, proofs, classification, formalization, and implementations. Besides mainstream research, ASHES puts some focus on new and emerging scenarios: This includes the Internet of Things (IoT), nuclear weapons inspections, arms control, consumer and infrastructure security, or supply chain security, among others. ASHES also welcomes works on special purpose hardware, such as lightweight, low-cost, and energy-efficient devices, or non-electronic security systems. Lejla Batina, Chip-Hong Chang, Ulrich Rührmair, Jakub Szefer |
CCS | 2 |
| 2024 | Steganographic Passport: An Owner and User Verifiable Credential for Deep Model IP Protection Without RetrainingabstractEnsuring the legal usage of deep models is crucial to promoting trustable, accountable, and responsible artifi-cial intelligence innovation. Current passport-based meth-ods that obfuscate model functionality for license-to-use and ownership verifications suffer from capacity and quality constraints, as they require retraining the owner model for new users. They are also vulnerable to advanced Expanded Residual Block ambiguity attacks. We propose Stegano-graphic Passport, which uses an invertible steganographic network to decouple license-to-use from ownership verification by hiding the user's identity images into the owner-side passport and recovering them from their respective user-side passports. An irreversible and collision-resistant hash function is used to avoid exposing the owner-side pass-port from the derived user-side passports and increase the uniqueness of the model signature. To safeguard both the passport and model's weights against advanced ambiguity attacks, an activation-level obfuscation is proposed for the verification branch of the owner's model. By jointly training the verification and deployment branches, their weights be-come tightly coupled. The proposed method supports agile licensing of deep models by providing a strong ownership proof and license accountability without requiring a sepa-rate model retraining for the admission of every new user. Experiment results show that our Steganographic Passport outperforms other passport-based deep model protection methods in robustness against various known attacks. Ruohan Meng, Chaohui Xu, Chip-Hong Chang |
CVPR | 4 |
| 2024 | Efficient Fast Additive Homomorphic Encryption Cryptoprocessor for Privacy-Preserving Federated Learning AggregationabstractPrivacy leakage is a critical concern of collaboratively training a large-scale deep learning model from multiple clients. To protect the local data, homomorphic encryption (e.g., Paillier) could be utilized for data aggregation on the central server. Nevertheless, even with CPU-optimized libraries or FPGA-based accelerators, the computing power and throughput limitations remain a stumbling block for practical deployment of Paillier scheme. In this paper, we present an efficient and high-throughput cryptoprocessor based on a recently introduced Fast Additive Homomorphic Encryption (FAHE) algorithm. For encryption, we incorporate the asymmetric decomposition, time multiplexing resource reuse and hard-macro based wide-bus logic operations to efficiently map the large (>40 kbits) integer multiplications for low latency FPGA implementation. For decryption, we propose a table lookup method for rapid modular reduction by leveraging the relative short modulus size of FAHE. The single large precomputed lookup table is carefully partitioned into multiple subtables and deployed in dual-port RAMs to enable resource-efficient parallel computation. The FAHE cryptoprocessor is implemented on a Xilinx ZCU102 FPGA board for performance evaluation and comparison. The results show that the throughput of our design is 354 × to 404 × higher than the state-of-the-art Paillier accelerators. Compared to the FAHE software implementation, the latency of our proposed design is 14.95 × and 11.42 × lower for encryption and decryption, respectively. Wenye Liu, Nazim Altar Koca, Chip-Hong Chang |
DATE | 3 |
| 2024 | Live Demonstration: Man-in-the-Middle Attack on Edge Artificial IntelligenceabstractDeep neural networks (DNNs) are susceptible to evasion attacks. However, digital adversarial examples are typically applied to pre-captured static images. The perturbations are generated by loss optimization with knowledge of target model hyperparameters and are added offline. Physical adversarial examples, on the other hand, tamper with the physical target or use a realistically fabricated target to fool the DNN. A sufficient number of pristine target samples captured under different varying environmental conditions are required to create the physical adversarial perturbations. Both digital and physical input evasion attacks are not robust against dynamic object-scene variations and the adversarial effects are often weak-ened by model reduction and quantization when the DNNs are implemented on edge artificial intelligence (AI) accelerator platforms. This demonstration presents a practical man-in-the-middle (MITM) attack on an edge DNN first reported in [1]. A tiny MIPI FPGA chip with hardened CSI-2 and D-PHY blocks is attached between the camera and the edge AI accelerator to inject unobtrusive stripes onto the RAW image data. The attack is less influenced by dynamic context variations such as changes in viewing angle, illumination, and distance of the target from the camera. Weiyang He, Wenye Liu, Chip-Hong Chang |
ISCAS | 5 |
| 2024 | Exploring Error Correction Circuits on RISC-V based Systems for Space ApplicationsabstractRISC-V systems are becoming increasingly adopted in space applications. SRAM (Static Random-Access Memory) data memory is a critical component that occupies a large portion of the processor peripheral system. SRAM is vulnerable to single-event upsets (SEUs). Existing studies mainly considered Hamming error correction codes (ECCs) for memory protection in RISC-V processor. In this paper, we explore different ECCs as well as the triple modular redundancy (TMR) as solutions to mitigate SEU effects on SRAMs for RISC-V system. To overcome the area and power overhead of TMR with higher fault tolerance than ECCs, we propose a dual modular redundancy (DMR) with ECC memory protection scheme. We conduct a comprehensive error analysis to evaluate different fault-tolerant designs under various radiation attack scenarios, utilizing real-world data to devise the fault-injection campaign. An efficient scheduling for the self-refresh operation of SRAMs is proposed to prevent the error accumulation. The proposed DMR with ECC design reduces the power and area overhead of TMR by 28% and 11% respectively and improve the error resilience of the SRAM significantly compared with Hsiao and Hamming ECC schemes. Nazim Altar Koca, Chip-Hong Chang, Anh-Tuan Do, Vishnu P. Nambiar |
ISCAS | 2 |
| 2024 | SPFL: A Self-Purified Federated Learning Method Against Poisoning AttacksabstractWhile Federated learning (FL) is attractive for pulling privacy-preserving distributed training data, the credibility of participating clients and non-inspectable data pose new security threats, of which poisoning attacks are particularly rampant and hard to defend without compromising privacy, performance or other desirable properties. In this paper, we propose a self-purified FL (SPFL) method that enables benign clients to exploit trusted historical features of locally purified model to supervise the training of aggregated model in each iteration. The purification is performed by an attention-guided self-knowledge distillation where the teacher and student models are optimized locally for task loss, distillation loss and attention loss simultaneously. SPFL imposes no restriction on the communication protocol and aggregator at the server. It can work in tandem with any existing secure aggregation algorithms and protocols for augmented security and privacy guarantee. We experimentally demonstrate that SPFL outperforms state-of-the-art FL defenses against poisoning attacks. The attack success rate of SPFL trained model remains the lowest among all defense methods in comparison, even if the poisoning attack is launched in every iteration with all but one malicious clients in the system. Meantime, it improves the model quality on normal inputs compared to FedAvg, either under attack or in the absence of an attack. Zizhen Liu, Weiyang He, Chip-Hong Chang, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | An Empirical Study of the Inherent Resistance of Knowledge Distillation Based Federated Learning to Targeted Poisoning AttacksabstractWhile the integration of Knowledge Distillation (KD) into Federated Learning (FL) has recently emerged as a promising solution to address the challenges of heterogeneity and communication efficiency, little is known about the security of these schemes against poisoning attacks prevalent in vanilla FL. From recent countermeasures built around KD, we conjecture that the way knowledge is distilled from the global model to the local models and the type of knowledge transfer by KD themselves offer some resilience against targeted poisoning attacks in FL. To attest this hypothesis, we systematize various adversary agnostic state-of-the-art KD-based FL algorithms for the evaluation of their resistance to different targeted poisoning attacks on two vision recognition tasks. Our empirical security-utility trade-off study indicates surprisingly good inherent immunity of certain KD-based FL algorithms that are not designed to mitigate these attacks. By probing into the causes of their robustness, the KD space exploration provides further insights into the balancing of security, privacy and efficiency triad in different FL settings. Weiyang He, Zizhen Liu, Chip-Hong Chang |
ATS | 3 |
| 2023 | ASHES '23: Workshop on Attacks and Solutions in Hardware SecurityabstractThe workshop on "Attacks and Solutions in HardwarE Security (ASHES)" welcomes any theoretical and practical works on hardware security, including attacks, solutions, countermeasures, proofs, classification, formalization, and implementations. Besides mainstream research, ASHES puts some focus on new and emerging scenarios: This includes the Internet of Things (IoT), nuclear weapons inspections, arms control, consumer and infrastructure security, or supply chain security, among others. ASHES also welcomes works on special purpose hardware, such as lightweight, low-cost, and energy-efficient devices, or non-electronic security systems. Lejla Batina, Chip-Hong Chang, Domenic Forte, Ulrich Rührmair |
CCS | 2 |
| 2023 | Energy-efficient NTT Design with One-bank SRAM and 2-D PE ArrayabstractIn Number Theoretic Transform (NTT) operation, more than half of the active energy consumption stems from memory accesses. Here, we propose a generalized design method to improve the energy efficiency of NTT operation by considering the effect of processing element (PE) geometry and memory organization on the data flow between PEs and memory. To decrease the number of data bits that are required to be accessed from the memory, a two-dimensional (2-D) PE array architecture is used. A pair of ping-pong buffers are proposed to transposed swap the coefficients to enable a single bank of memory to be used with the 2-D PE array to reduce the average memory bit access energy without compromising the throughput. Our experimental results show that this design method can produce NTT accelerators with up to 69.8% saving in average energy consumption compared with the existing designs based on multi-bank SRAM and one-bank SRAM with one-dimensional PE array with the same number of PEs and total memory size. Jianan Mu, Huajie Tan, Haotian Lu 0002, Chip-Hong Chang, Shengwen Liang, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
DATE | 5 |
| 2023 | White-Box Adversarial Attacks on Deep Learning-Based Radio Frequency Fingerprint IdentificationabstractRadio frequency fingerprint identification (RFFI) is an emerging technique for the lightweight authentication of wireless Internet of things (IoT) devices. RFFI exploits unique hardware impairments as device identifiers, and deep learning is widely deployed as the feature extractor and classifier for RFFI. However, deep learning is vulnerable to adversarial attacks, where adversarial examples are generated by adding perturbation to clean data for causing the classifier to make wrong predictions. Deep learning-based RFFI has been shown to be vulnerable to such attacks, however, there is currently no exploration of effective adversarial attacks against a diversity of RFFI classifiers. In this paper, we report on investigations into white-box attacks (non-targeted and targeted) using two approaches, namely the fast gradient sign method (FGSM) and projected gradient descent (PGD). A LoRa testbed was built and real datasets were collected. These adversarial examples have been experimentally demonstrated to be effective against convolutional neural networks (CNNs), long short-term memory (LSTM) networks, and gated recurrent units (GRU). Junqing Zhang, Guanxiong Shen, Alan Marshall 0001, Chip-Hong Chang |
ICC | 5 |
| 2023 | Hardware-efficient Softmax Approximation for Self-Attention NetworksabstractSelf-attention networks such as Transformer have become state-of-the-art models for natural language processing (NLP) problems. Softmax function, which serves as a normalizer to produce attention scores, turns out to be a severe throughput and latency bottleneck of a Transformer network. Softmax datapath consists of data-dependent sequential nonlinear exponentiation and division operations, which are not amenable to pipelining and parallelism, nor can they be directly linearized for pretrained models without substantial accuracy drop. In this paper, we proposed a hardware efficient Softmax approximation which can be used as a direct plug-in substitution into pretrained transformer network to accelerate NLP tasks without compromising its accuracy. Experiment results on FPGA implementation show that our design outperforms vanilla Softmax designed using Xilinx IPs with 15x less LUTs, 55x less registers and 23x lower latency at similar clock frequency and less than 1% accuracy drop on main language benchmark tasks. We also propose a pruning method to reduce the input entropy of Softmax for NLP problems with high number of inputs. It was validated on CoLA task to achieve a further 25% reduction of latency. Nazim Altar Koca, Anh-Tuan Do, Chip-Hong Chang |
ISCAS | 3 |
| 2023 | A Lightweight PUF-based Secure Group Key Agreement Protocol for Wireless Sensor NetworksabstractWireless sensor networks (WSNs) have gained considerable popularity in a wide range of applications such as military, healthcare, transportations, and environmental sensing. Data produced in these applications is highly sensitive and requires a high level of security protection. However, individual nodes in WSNs are typically resource constrained and have limited computing power to protect memory stored secrets, which are vulnerable to tampering and physical probing. As group messaging is commonly used in WSNs for efficient message exchanges among sensor nodes, this paper presents a lightweight secure group key agreement protocol using Physical Unclonable Function (PUF) as a hardware root of trust. The proposed scheme establishes secure group authentication and group session key simultaneously for all participating members of the group without resorting to complex public-key algorithms. By hiding the prover's authentication secrets in a secure mask, the verifier does not have to store the secrets but recover them for authentication by querying its PUF. The proposed protocol enables lightweight cluster head authentication at the sensor node and prevents stolen-verifier attack at the cluster head. Besides, it is robust against memory probing attacks at all group devices and man-in-the-middle attacks on the communication channel. Among existing PUF-based group key establishment protocols, it requires zero secret storage cost and exhibits excellent overall computation and communication performance. Wenye Liu, Chip-Hong Chang |
ISCAS | 3 |
| 2023 | Scalable and Conflict-Free NTT Hardware Accelerator Design: Methodology, Proof, and ImplementationabstractNumber theoretic transform (NTT) is useful for the acceleration of polynomial multiplication, which is the main performance bottleneck in the next-generation cryptographic schemes. Different NTT-based cryptographic algorithms have different security settings. The diverse application scenarios introduce different cost-performance tradeoffs and hardware constraints. Motivated by the emerging demand for more versatile NTT hardware accelerators, we propose a new design methodology that can generate area-efficient and high-performance NTT accelerators for any length and modulus of NTT polynomials and single processing element (PE) or PE array with a varying number of layers. The proposed NTT accelerator architecture pivots on a conflict-free memory access pattern for adaptation to different combinations of security and PE array configuration parameters. The proposed memory access pattern is formally proved to be conflict-free for any parametric configurations. The criterion for read-after-write conflict without pipeline stall is also established. Our proposed design methodology can produce NTT accelerators with single PE or multilayer PE array for different polynomial size and modulus, with hardware area and computational efficiency comparable to accelerators customized for a fixed set of parameters. Our proposed methodology produces parameterized accelerator with higher scalability than the existing parameterized accelerator design. On average, the accelerators generated by our proposed method are 71.4% more area-time efficient. Up to 30.7% area-time reduction over the most area-time efficient state-of-the-art scalable NTT accelerator can be achieved for the same security parameters. Jianan Mu, Wen Wang 0007, Yizhong Hu, Chip-Hong Chang, Junfeng Fan, Jing Ye 0001, Yuan Cao 0003, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | A New Reconfigurable True Random Number Generator and Physical Unclonable Function Unified Chip With On-Chip Auto-CalibrationabstractTrue random number generator (TRNG) and physical unclonable function (PUF) have been extensively used to secure low-cost Internet of Things (IoT) endpoints. In this paper, a lightweight reconfigurable TRNG and PUF unified design for custom chip implementation is proposed. The reconfigurable structure consists of a pair of ring oscillators (ROs) with interposed multi-way switches for RO length reconfiguration and shared counters for on-chip calibration. Jitter noise of ROs and metastability of arbiter are harmonized for TRNG operation, while process variations of ROs are extracted for PUF operation. The conflicting requirements on frequency deviation for the randomness of TRNG and the reliability of PUF are resolved by an on-chip calibrator, which automatically selects and stores a challenge with a small frequency difference in TRNG mode upon manufacturing and masks unreliable challenges with large frequency difference during PUF enrollment. Leveraging the advantage of custom chip design, the basic delay cell of the reconfigurable ROs is realized by current starved inverter in weak inversion to minimize the power consumption, increase the jitter, and avail its larger process variation. A new lightweight secure mutual authentication protocol is also proposed to effectively thwart machine learning, replay and man-in-the-middle attacks using only the underlying TRNG and PUF without requiring any other security primitives. The proposed TRNG-PUF design is prototyped with a standard 40 nm 1.1 V CMOS process. It occupies a small footprint of 24,$316~\pmb {\mu m^{2}}$. Measured results of the packaged chips show an average energy efficiency of 7.42 pJ/bit in TRNG operation and 0.10 pJ/bit in PUF operation. The bitstreams generated by the test chips passed NIST SP 800-22 and 90B tests, autocorrelation test, and FFT test. Yuan Cao 0003, Wanyi Liu, Jing Ye 0001, Chip-Hong Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | Guest Editorial Special Issue on the Asian Hardware Oriented Security and Trust Symposium (AsianHOST 2022)abstractAsian Hardware Oriented Security and Trust Symposium (AsianHOST) is an annual symposium that aims to facilitate the rapid growth of hardware-based security research and development. Hardware security is a fashionable research area in both industry and academia. Its scope is consistently growing to embrace secure design, manufacturing, and deployment of modern and emerging interoperable computing, communication, storage devices, and circuits and systems. The 7th Asian Hardware Oriented Security and Trust Symposium (AsianHOST 2022) was held in hybrid mode on December 14–16 in Singapore. Among all the accepted contributions presented at the conference, a subset of top-rated articles was selected and invited for this Special Issue. The invited articles included extended new technical contributions and results and went through a peer-review process consisting of expert reviewers in the related topics. A brief description of the selected articles is as follows. Chip-Hong Chang, Pingqiang Zhou, Yuan Cao 0003, Qiang Liu 0011 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | An Imperceptible Data Augmentation Based Blackbox Clean-Label Backdoor Attack on Deep Neural NetworksabstractDeep neural networks (DNNs) have permeated into many diverse application domains, making them attractive targets of malicious attacks. DNNs are particularly susceptible to data poisoning attacks. Such attacks can be made more venomous and harder to detect by poisoning the training samples without changing their ground-truth labels. Despite its pragmatism, the clean-label requirement imposes a stiff restriction and strong conflict in simultaneous optimization of attack stealth, success rate, and utility of the poisoned model. Attempts to circumvent the pitfalls often lead to a high injection rate, ineffective embedded backdoors, unnatural triggers, low transferability, and/or poor robustness. In this paper, we overcome these constraints by amalgamating different data augmentation techniques for the backdoor trigger. The spatial intensities of the augmentation methods are iteratively adjusted by interpolating the clean sample and its augmented version according to their tolerance to perceptual loss and augmented feature saliency to target class activation. Our proposed attack is comprehensively evaluated on different network models and datasets. Compared with state-of-the-art clean-label backdoor attacks, it has lower injection rate, stealthier poisoned samples, higher attack success rate, and greater backdoor mitigation resistance while preserving high benign accuracy. Similar attack success rates are also demonstrated on the Intel Neural Compute Stick 2 edge AI device implementation of the poisoned model after weight-pruning and quantization. Chaohui Xu, Wenye Liu, Chip-Hong Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | PUF-Based Mutual Authentication and Key Exchange Protocol for Peer-to-Peer IoT ApplicationsabstractPeer to Peer (P2P) or direct connection IoT has become increasingly popular owing to its lower latency and higher privacy compared to database-driven or server-based IoT. However, wireless vulnerabilities raise severe concerns on IoT device-to-device communication. This is further aggravated by the challenge to achieve lightweight direct mutual authentication and secure key exchange between IoT peer nodes in P2P IoT applications. Physical unclonable function (PUF) is a key enabler to lightweight, low-power and secure authentication of resource-constrained devices in IoT. Nevertheless, current PUF-enabled authentication protocols, with or without the challenge-response pairs (CRPs) of each of its interlocutors stored in the verifier's side, are incompatible for P2P IoT scenarios due to the security, storage and computing power limitations of IoT devices. To solve this problem, a new lightweight PUF-based mutual authentication and key exchange protocol is proposed. It allows two resource-constrained PUF embedded endpoint devices to authenticate each other directly without the need for local storage of CRPs or any private secrets, and simultaneously establish the session key for secure data exchange without resorting to the public-key algorithm. The proposed protocol is evaluated using the game-based formal security analysis method as well as the automatic security analysis tool ProVerif to corroborate its mutual authenticity, secrecy, and resistance against replay and man-in-the-middle (MITM) attacks. Using two Avnet Ultra96-V2 boards to emulate the two IoT endpoint devices, a physical prototype system is also constructed to demonstrate and validate the feasibility of the proposed secure P2P connection scheme. A comparative analysis shows that the proposed protocol outperforms related protocols in terms of security features, computational complexity as well as communication and storage costs. Wenye Liu, Chongyan Gu, Chip-Hong Chang |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | ASHES 2022 - 6th Workshop on Attacks and Solutions in Hardware SecurityabstractThe workshop on "Attacks and Solutions in HardwarE Security (ASHES)" welcomes any theoretical and practical works on hardware security, including attacks, solutions, countermeasures, proofs, classification, formalization, and implementations. Besides mainstream research, ASHES puts some focus on new and emerging scenarios: This includes the Internet of Things (IoT), nuclear weapons inspections, arms control, consumer and infrastructure security, or supply chain security, among others. ASHES also welcomes dedicated works on special purpose hardware, such as lightweight, low-cost, and energy-efficient devices, or non-electronic security systems. The workshop hosts four different paper categories: Apart from regular and short papers, this includes works that systematize and structure a certain (sub-)area (so-called "Systematization of Knowledge" (SoK) papers), and so-termed "Wild-and-Crazy" (WaC) papers, which distribute seminal ideas at an early conceptual stage. This summary gives a brief overview of the sixth edition of the workshop, which took place virtually on November 11, 2022 in Los Angeles, California, USA, as a post-conference satellite workshop of ACM CCS. Chip-Hong Chang, Domenic Forte, Debdeep Mukhopadhyay, Ulrich Rührmair |
CCS | 1 |
| 2022 | Area, Time and Energy Efficient Multicore Hardware Accelerators for Extended Merkle Signature SchemeabstractThis paper addresses a barrier that prevents the timely adoption of post-quantum signature algorithms, such as the eXtended Merkle Signature Scheme (XMSS), due to its lack of fast, cost-effective and energy-efficient hardware accelerators. Two new architectures that use more than one hash core are proposed for the first time to significantly reduce the latency of two bottleneck XMSS operations, namely key generation and signature generation, for which the speed of existing hardware accelerators is still apparently inadequate. The first proposed multi-core design uses block RAM and a simplified data flow to maximize the use of$p$hash cores concurrently in three major sequential stages of computation, i. e., Winternitz One-time Signature (WOTS), L-tree and Merkle tree. The second proposed multi-core design adds a dedicated hash core for tree hashing in the L-tree and Merkle tree while keeping the$p$hash cores solely for chain hashing in WOTS. The dedicated hash core leapfrogs between the L-tree and Merkle tree and computes concurrently with the$p$hash cores to keep the$p+1$hash cores active most of the time while minimizing the storage requirement and energy consumption. Both designs are implemented on a 28 nm ATRIX-7 FPGA chip. Experimental results show that both proposed accelerators with$p=8$operate at a much faster speed and consume significantly less hardware resources and energy than all existing XMSS accelerators. Specifically, they are$\sim 8\times $and$\sim 6\times $faster than the fastest reported design in key generation and signature generation operations, respectively. Yuan Cao 0003, Yanze Wu, Lan Qin, Chip-Hong Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | An Efficient Full Hardware Implementation of Extended Merkle Signature SchemeabstractThis paper presents a full hardware implementation of the eXtended Merkle Signature Scheme (XMSS), a NIST approved and IETF RFC specified post-quantum cryptography (PQC) algorithm. An optimized node traversal is proposed to enable efficient memory utilization without compromising the computational latency of the L-tree and Merkle tree construction, which are two key components used for the compression of the Winternitz One-Time Signature (WOTS) public key in XMSS. The computation of the authentication path during signature generation has also been significantly sped up by our proposed hardware implementation of the Buchmann, Dahmen, and Schneider (BDS) algorithm. Our implementation has completely avoided the use of block random-access memory, which is known to be vulnerable to side-channel attacks. The memory requirement has been highly optimized for implementation with small flip-flop chains and register counters as pointers for fast data access. To the best of our knowledge, this is the first full hardware implementation of all threekey generation,signingandverificationoperations of XMSS. The design has been prototyped and evaluated on a 28 nm FPGA platform to demonstrate its performance improvements over the most efficient software and hardware/software co-design methods reported to date. Specifically, it increases the computational efficiency of the best reported XMSS implementation forkey generationandsignature generationby about 20% and 50%, respectively. It can also run at 10% higher clock speed than the fastest hardware implementation ofsignature verificationin FPGA with 8% lower hardware resource utilization. Yuan Cao 0003, Yanze Wu, Wen Wang 0007, Jing Ye 0001, Chip-Hong Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | A New Energy-Efficient and High Throughput Two-Phase Multi-Bit per Cycle Ring Oscillator-Based True Random Number GeneratorabstractOscillator-based elementary true random number generator (TRNG) uses a slow jittery ring oscillator (RO) to sample a fast RO. The ROs are always on but most of the oscillatory cycles of the fast RO are not sampled into random bits. In this paper, a new lightweight TRNG design is proposed to minimize the power wasted by the superfluous oscillations. Random bits are extracted from both phases of the slow ROs to increase the throughput and the fast RO is activated only during the narrow transition time difference between two symmetrically designed slow ROs. The slow jittery ROs are implemented using current starved inverters biased in the weak inversion region to reduce their power consumption. Their jitter amplitudes are increased by lowering the oscillation frequency and reducing the drain current of the transistors. The narrow jittery pulse generated by the differential pair of slow ROs is quantized by the fastest three-stage RO. Two random bits from each phase of the jittery ROs can be extracted by using a gigahertz dynamic toggled D flip-flop counter to count the number of oscillatory cycles of the fast RO. The proposed TRNG is fabricated in a standard 65 nm 1.2 V CMOS process. Measurement results of the fabricated chips show that the proposed TRNG consumes merely$260~\mu \text{W}$at a bit rate of 52 Mbps. It outperforms the state-of-art on-chip jitter-based TRNGs with the best figure-of-merit of 5 pJ/bit and the smallest footprint of$366~\mu \text{m}^{2}$. Its generated bit sequence passes the statistical randomness tests including National Institute of Standards and Technology (NIST) test, Auto Correlation Factor (ACF) test and bias. The mean redundancy of the ten tested chips is measured to be less than 10−5bit/symbol. Yuan Cao 0003, Xiaojin Zhao, Wenhan Zheng, Chip-Hong Chang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | A DNN Fingerprint for Non-Repudiable Model Ownership Identification and Piracy DetectionabstractA high-performance Deep Neural Network (DNN) model is a valuable intellectual property (IP) since designing and training such a model from scratch is very costly. Model transfer learning, compression and retraining are commonly used by pirates to evade detection or even redeploy the pirated models for new applications without compromising performance. This paper presents a novel non-intrusive DNN IP fingerprinting method that can detect pirated models and provide a non-repudiable and irrevocable ownership proof simultaneously. The fingerprint is derived from projecting a subset of front-layer weights onto a model owner identity defined random space to enable a distinguisher to differentiate pirated models that are used in the same application or retrained for a different task from originally designed DNN models. The proposed method generates compact and irrevocable fingerprints against model IP misappropriation and ownership fraud. It requires no retraining and makes no modification to the original model. The proposed fingerprinting method is evaluated on nine original DNN models trained on CIFAR-10, CIFAR-100, and ImageNet-10. It is demonstrated to have the highest discriminative power among existing fingerprinting methods in detecting pirated models deployed for the same and different applications, and fraudulent model IP ownership claims. Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Live Demonstration: Event-Driven Physical Unclonable Function for Proactive Monitoring System by Dynamic Vision SensorabstractDynamic Vision Sensor (DVS) is a representative neuromorphic camera that responds to changes in luminosity due to movement in the scene. A fingerprinting scheme based on physical unclonable functions (PUF) is implemented to bind a DVS footage to its camera identity. This process is only triggered by incidents of interest to minimise computation. A transmission protocol based on the PUF is designed, alerting the server to any tampering attempts. In this demonstration, we use a DVS to capture a speeding incident simulated by a remote-controlled (RC) car. The baseline PUF database is obtained through a postlayout Monte Carlo simulation of the event-driven PUF. It will be embedded as a security provider in a Raspberry Pi 4 (RPi4), that serves as the DVS embedded system. Man-in-the-middle attacks will be simulated against the system. Jian Xian Soo, Chip-Hong Chang |
ISCAS | 3 |
| 2021 | Fingerprinting Deep Neural Networks - a DeepFool ApproachabstractA well-trained deep learning classifier is an expensive intellectual property of the model owner. However, recently proposed model extraction attacks and reverse engineering techniques make model theft possible and similar quality deep learning solution reproducible at a low cost. To protect the interest and revenue of the model owner, watermarking on Deep Neural Network (DNN) has been proposed. However, the extra components and computations due to the embedded watermark tend to interfere with the model training process and result in inevitable degradation in classification accuracy. In this paper, we utilize the geometry characteristics inherited in the DeepFool algorithm to extract data points near the classification boundary of the target model for ownership verification. As the fingerprint is extracted after the training process has been completed, the original achievable classification accuracy will not be compromised. This countermeasure is founded on the hypothesis that different models possess different classification boundaries determined solely by the hyperparameters of the DNN and the training it has undergone. Therefore, given a set of fingerprint data points, a pirated model or its post-processed version will produce similar prediction but another originally designed and trained DNN for the same task will produce very different prediction even if they have similar or better classification accuracy. The effectiveness of the proposed Intellectual Property (IP) protection method is validated on the CIFAR-10, CIFAR-100 and ImageNet datasets. The results show a detection rate of 100% and a false positive rate of 0% for each dataset. More importantly, the fingerprint extraction and its run time are both dataset independent. It is on average ~130× faster than two state-of-the-art fingerprinting methods. Chip-Hong Chang |
ISCAS | 2 |
| 2021 | Secure Mutual Authentication and Key-Exchange Protocol between PUF-Embedded IoT EndpointsabstractDevice authentication and key exchange protocol are essential front line of access controls in IoT security, and physical unclonable function (PUF) is a key enabler to lightweight, low-power and secure authentication of internet enabled endpoint devices in IoT. Current PUF-enabled authentication protocol requires the verifier to store a sufficiently large number of challenge-response pairs (CRPs) of each of its interlocutors, which makes the protocol impractical in application scenarios where the verifier is a resource-constrained device, especially when the verifier needs to communicate with multiple PUF- embedded endpoints. To solve this problem, a new lightweight PUF-based mutual authentication and key-exchange protocol is proposed in this paper to allow two resource-constrained PUF embedded endpoint devices to authenticate each other without the need to store the CRPs locally, and simultaneously establish the session key for secure data exchange without resorting to public-key algorithm. The PUF response reliability is mitigated by the use of reverse fuzzy extractor to offload the compute- intensive error decoding to the server during the initialization phase of the protocol. The proposed protocol is evaluated using ProVerif to corroborate its secrecy, mutual authenticity as well as resistance against replay and man-in-the-middle attacks. Chip-Hong Chang |
ISCAS | 2 |
| 2021 | A Modeling Attack Resistant Deception Technique for Securing Lightweight-PUF-Based AuthenticationabstractSilicon physical unclonable function (PUF) has emerged as a promising spoof-proof solution for low-cost device authentication. Due to practical constraints in preventing phishing through a public network or insecure communication channels, simple PUF-based authentication protocol with unrestricted queries and transparent responses is vulnerable to modeling and replay attacks. In this article, we present a modeling attack resistant PUF-based mutual authentication scheme to mitigate the practical limitations in applications where a resource-rich server authenticates a device with no strong restriction imposed on the type of PUF design or any additional protection on the binary channel used for the authentication. Our scheme uses an active deception protocol to prevent machine learning (ML) attacks on a device with a monolithic integration of a genuine strong PUF (SPUF), a fake PUF, a pseudorandom number generator (PRNG), a register, a binary counter, a comparator, and a simple controller. The hardware encapsulation makes the collection of challenge-response pairs (CRPs) easy for model building during enrollment but prohibitively time consuming upon device deployment through the same interface. A genuine server can perform a mutual authentication with the device using a combined fresh challenge contributed by both the server and the device. The message exchanged in clear cannot be manipulated by the adversary to derive unused authentic CRPs. The adversary will have to either wait for an impractically long time to collect enough real CRPs by directly querying the device or the ML model derived from the collected CRPs will be poisoned. The false PUF multiplexing is fortified against the prediction of waiting time by doubling the time penalty for every unsuccessful guess. Our implementation results on field-programmable gate array (FPGA) device and security analysis have corroborated the low hardware overheads and attack resistance of the proposed deception protocol. Chongyan Gu, Chip-Hong Chang, Weiqiang Liu 0001, Shichao Yu, Yale Wang, Máire O'Neill |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | An Overview of Hardware Security and Trust: Threats, Countermeasures, and Design ToolsabstractHardware security and trust have become a pressing issue during the last two decades due to the globalization of the semiconductor supply chain and ubiquitous network connection of computing devices. Computing hardware is now an attractive attack surface for launching powerful cross-layer security attacks, allowing attackers to infer secret information, hijack control flow, compromise system root-of-trust, steal intellectual property (IP), and fool machine learners. On the other hand, security practitioners have been making tremendous efforts in developing protection techniques and design tools to detect hardware vulnerabilities and fortify hardware design against various known hardware attacks. This article presents an overview of hardware security and trust from the perspectives of threats, countermeasures, and design tools. By introducing the most recent advances in hardware security research and developments, we aim to motivate hardware designers and electronic design automation tool developers to consider the new challenges and opportunities of incorporating an additional dimension of security into robust hardware design, testing, and verification. Wei Hu 0008, Chip-Hong Chang, Anirban Sengupta 0003, Swarup Bhunia, Ryan Kastner, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Identification of FSM State Registers by Analytics of Scan-Dump DataabstractBig data analytics have gained tremendous successes in mining valuable information in various fields. However, its potential to solve complex problems in hardware security has not been adequately tapped. This paper presents a non-invasive approach to identify the state registers of a finite state machine (FSM) in an integrated chip. The state registers of the FSM are mined from the scan-dump data by exploiting the strongly connected property and chronologically correlated state codes of the FSM. The sequence of data scanned out of each scan register is partitioned into non-overlapping strings of high weighted frequencies by a string-matching algorithm. A coherency between a pair of registers is defined and computed based on the partitioned strings. The dimension of the coherency matrix is first reduced by pruning some registers of low influence by a regression analysis. The registers are then clustered to minimize the within-cluster variances based on their coherency values. The proposed scheme is applied to some IP cores from OpenCores. The experimental results show that our scheme can correctly identify the FSM state registers in most designs with high hit rate. Aijiao Cui, Chengkang He, Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Stealthy and Robust Glitch Injection Attack on Deep Learning Accelerator for Target With Variational ViewpointabstractDeep neural network (DNN) accelerators overcome the power and memory walls for executing neural-net models locally on edge-computing devices to support sophisticated AI applications. The advocacy of “model once, run optimized anywhere” paradigm introduces potential new security threat to edge intelligence that is methodologically different from the well-known adversarial examples. Existing adversarial examples modify the input samples presented to an AI application either digitally or physically to cause a misclassification. Nevertheless, these input-based perturbations are not robust or surreptitious on multi-view target. To generate a good adversarial example for misclassifying a real-world target of variational viewing angle, lighting and distance, a decent number of target’s samples are required to extract the rare anomalies that can cross the decision boundary. The feasible perturbations are substantial and visually perceptible. In this paper, we propose a new glitch injection attack on DNN accelerator that is capable of misclassifying a target under variational viewpoints. The glitches injected into the computation clock signal induce transitory but disruptive errors in the intermediate results of the multiply-and-accumulate (MAC) operations. The attack pattern for each target of interest consists of sparse instantaneous glitches, which can be derived from just one sample of the target. Two modes of attack patterns are derived, and their effectiveness are demonstrated on four representative ImageNet models implemented on the Deep-learning Processing Unit (DPU) of FPGA edge and its DNN development toolchain. The attack success rates are evaluated on 118 objects in 61 diverse sensing conditions, including 25 viewing angles (−60° to 60°), 24 illumination directions and 12 color temperatures. In the covert mode, the success rates of our attack exceed existing stealthy adversarial examples by more than 16.3%, with only two glitches injected into ten thousands to a million cycles for one complete inference. In the robust mode, the attack success rates on all four DNNs are more than 96.2% with an average glitch intensity of 1.4% and a maximum glitch intensity of 10.2%. Wenye Liu, Chip-Hong Chang, Fan Zhang 0010 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | ASHES 2020: 4th Workshop on Attacks and Solutions in Hardware Security
Chip-Hong Chang, Stefan Katzenbeisser 0001, Ulrich Rührmair, Patrick Schaumont |
CCS | 1 |
| 2020 | Imperceptible Misclassification Attack on Deep Learning Accelerator by Glitch InjectionabstractThe convergence of edge computing and deep learning empowers endpoint hardwares or edge devices to perform inferences locally with the help of deep neural network (DNN) accelerator. This trend of edge intelligence invites new attack vectors, which are methodologically different from the well-known software oriented deep learning attacks like the input of adversarial examples. Current studies of threats on DNN hardware focus mainly on model parameters interpolation. Such kind of manipulation is not stealthy as it will leave non-erasable traces or create conspicuous output patterns. In this paper, we present and investigate an imperceptible misclassification attack on DNN hardware by introducing infrequent instantaneous glitches into the clock signal. Comparing with falsifying model parameters by permanent faults, corruption of targeted intermediate results of convolution layer(s) by disrupting associated computations intermittently leaves no trace. We demonstrated our attack on nine state-of-the-art ImageNet models running on Xilinx FPGA based deep learning accelerator. With no knowledge about the models, our attack can achieve over 98% misclassification on 8 out of 9 models with only 10% glitches launched into the computation clock cycles. Given the model details and inputs, all the test images applied to ResNet50 can be successfully misclassified with no more than 1.7% glitch injection. Wenye Liu, Chip-Hong Chang, Fan Zhang 0010, Xiaoxuan Lou |
DAC | 2 |
| 2020 | Reducing Temperature Induced Unreliability in Sub-Threshold Strong PUFs through Circuit ModelingabstractDue to increased adoption of IoT devices, Physically Unclonable Function (PUF) circuits are essential as a lightweight security primitive. However existing PUFs suffer from imperfect reliability due to environmental variation. In this paper, we propose a PUF-server paradigm in which the server hosts a temperature model of the PUF, and the PUF informs the server of its response, and temperature through an integrated sensor. The model can predict responses for two different strong PUFs across -45°C to 90°C and thus avoids losing challenge-response pairs(CRPs) owing to low mismatch. Simulations show the model-based reliability improves to 97.9% - 99.5% compared with conventional reliability of 83.9% - 86.3%. Conversely, a loss of 25% - 40% of the CRPs is observed in order to improve the conventional reliability to match the model-based reliability. These results are also shown to be robust against quantization errors of a temperature sensor. Nimesh Shah, Sumon Kumar Bose, Chip-Hong Chang, Arindam Basu |
ISCAS | 3 |
| 2020 | Fired Neuron Rate Based Decision Tree for Detection of Adversarial Examples in DNNsabstractDeep neural network (DNN) is a prevalent machine learning solution to computer vision problems. The most criticized vulnerability of deep learning is its susceptibility towards adversarial images crafted by maliciously adding infinitesimal distortions to the benign inputs. Such negatives can fool a classifier. Existing countermeasures against these adversarial attacks are mainly developed based on software model of DNNs by using modified training during learning or modified input during testing, modifying networks or changing loss/activation functions, or relying on add-on models for classifying unseen examples. These approaches do not consider the optimization for hardware implementation of the learning models. In this paper, a new thresholding method is proposed based on comparators integrated into the most discriminative layers of the DNN determined by their layer-wise fired neuron rates between adversarial and normal inputs. Effectiveness of the method is validated on the ImageNet dataset with 8-bit truncated models for the state-of-the-art DNN architectures. A high detection rate of up to 98% with only 4.5% of false positive rate is achieved. The results show a significant improvement on both detection rate and false positive rate compared with previous countermeasures against the most practical non-invasive universal perturbation attack on deep learning based AI chip. Wenye Liu, Chip-Hong Chang |
ISCAS | 3 |
| 2020 | A PUF-Based Data-Device Hash for Tampered Image Detection and Source Camera IdentificationabstractWith the increasing prevalent of digital devices and their abuse for digital content creation, forgeries of digital images and video footage are more rampant than ever. Digital forensics is challenged into seeking advanced technologies for forgery content detection and acquisition device identification. Unfortunately, existing solutions that address image tampering problems fail to identify the device that produces the images or footage while techniques that can identify the camera is incapable of locating the tampered content of its captured images. In this paper, a new perceptual data-device hash is proposed to locate maliciously tampered image regions and identify the source camera of the received image data as a non-repudiable attestation in digital forensics. The presented image may have been either tampered or gone through benign content preserving geometric transforms or image processing operations. The proposed image hash is generated by projecting the invariant image features into a physical unclonable function (PUF)-defined Bernoulli random space. The tamper-resistant random PUF response is unique for each camera and can only be generated upon triggered by a challenge, which is provided by the image acquisition timestamp. The proposed hash is evaluated on the modified CASIA database and CMOS image sensor-based PUF simulated using 180 nm TSMC technology. It achieves a high tamper detection rate of 95.42% with the regions of tampered content successfully located, a good authentication performance of above 98.5% against standard content-preserving manipulations, and 96.25% and 90.42%, respectively, for the more challenging geometric transformations of rotation (0 ~ 360°) and scaling (scale factor in each dimension: 0.5). It is demonstrated to be able to identify the source camera with 100% accuracy and is secure against attacks on PUF. Yuan Cao 0003, Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2020 | Ed-PUF: Event-Driven Physical Unclonable Function for Camera Authentication in Reactive Monitoring SystemabstractAs surveillance footage plays an increasingly significant role in law enforcement, it is imperative to ensure the integrity of recorded video data and the authenticity of its originator, and instill situation awareness into these monitoring systems with a fidelity record of the incidents. Unfortunately, existing frame-based networked surveillance systems could only partially fulfill these requirements. The emerging Dynamic Vision Sensor (DVS) sheds new light on solving this problem with its completely different sensor design, i.e., DVS responds only to temporal intensity change and records only sparse asynchronous address-events with precise timing information. Motivated by the reduced data size of activities and the prevention of privacy intrusion of subjects under surveillance as well as other appealing attributes, this work introduces the first event-driven physical unclonable function (Ed-PUF) system to fill the forensic gap of simultaneously authenticating the event data integrity and source camera identity for reactive monitoring by DVS camera. New DVS sensor architecture is proposed with negligible modifications made to the original DVS pixel. The Ed-PUF response bit can only be triggered by and uniquely dependent on the asynchronous addressed event without being interfered by the simultaneous firing of other address events. Address event streams are securely transmitted with an event package tag created by a keyed hash-based message authentication code with the key being the Ed-PUF response. A secure protocol to authenticate the identity of DVS camera and the integrity of address events transmitted through cellular network is also proposed. A camera lock is embedded to protect against severing and splicing the inter-chip connectivity within the camera for raw PUF responses. The proposed system is evaluated using raw PUF data obtained by post-layout Monte Carlo simulation in UMC 180nm technology and real event stream captured by a DVS camera. The proposed Ed-PUF has been demonstrated to have excellent uniqueness, randomness and reliability. Collision test is also conducted to show that the quality of DVS imaging is not compromised. Besides keeping the hardware/power/timing overheads low, the proposed scheme is also analyzed to be resilient against multiple attack scenarios. Xiaojin Zhao, Takashi Sato 0001, Yuan Cao 0003, Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | A New Polarization Image Demosaicking Algorithm by Exploiting Inter-Channel Correlations With Guided FilteringabstractThis paper presents a fast and effective polarization image demosaicking algorithm, which explores inter-channel dependency of Stokes parameters for the minimization of residual aliasing artifacts after cubic spline interpolation. A guided filtering approach is used for denoising. An optimization based on the confidence level of the aforementioned guided filtering, the correlations between the demosaicked image and input, as well as the total intensity, angle and degree of linear polarization, is constructed and solved with Newton's method. Experimental results demonstrate that the proposed algorithm can surpass the existing methods in terms of both objective root mean squared error and structural similarity index by at least 36.0% and 3.4%, respectively, and by close visual inspection of the clarity of objects in the angle and degree of linear polarization images. The proposed algorithm consists of only convolutions and element-wise operations, making it fast and parallelizable for efficient GPU acceleration. An image of size 512 × 612 × 4 can be processed within 10 s on i7-6700k CPU, and gains further 5 times speedup with M4000M GPU. ShuMin Liu, Jiajia Chen 0002, Yuan Xun, Xiaojin Zhao, Chip-Hong Chang |
IEEE Trans. Image Process. | 5 |
| 2020 | A 1036-F2/Bit High Reliability Temperature Compensated Cross-Coupled Comparator-Based PUFabstractIn this article, a compact physical unclonable function (PUF) based on cross-coupled comparator is presented. Featuring a positive feedback response generation mechanism, the mismatch in analog signals between the cross-coupled transistor pair is quickly amplified to prevent its polarity from flipping by the temporal noise. The rapid enlargement of noise margin by the sense amplifier also contributes to stabilizing the response against supply voltage variations. To improve its temperature stability, the counteracting effect of complementary-to-absolutetemperature (CTAT) and proportional-to-absolute-temperature (PTAT) drives are considered in sizing the bit cell transistors. The proposed design is fabricated in a standard 65-nm CMOS process. The bit cell occupies an area of only 4.38 μm2(i.e., 1036 F2), and the overall PUF chip consumes 2.98 pJ/bit at the throughput of 8 Mb/s, of which only 1.61 pJ/bit is due to the PUF's core. With the uniqueness measured to be 49.53%, the unpredictability of the fabricated PUF chips is validated by autocorrelation function and NIST randomness tests. Compared with the state-of-the-art implementations, the proposed PUF has the lowest native response instability of 1.46% with 500 repeated PUF readouts at 27 °C and 1.2 V. By varying the operating temperature from -50 °C to 150 °C in a step size of 10 °C and the supply voltage from 1.0 to 1.4 V in a step size of 0.1 V simultaneously, the average reliability of the proposed PUF obtained from the 2-D plot of all operating conditions is found to be 96.87% without correction and 99.31% with spatial majority voting (SMV). Yiheng Wu, Xiaojin Zhao, Yuan Cao 0003, Chip-Hong Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | ASHES 2019: 3rd Workshop on Attacks and Solutions in Hardware Security
Chip-Hong Chang, Daniel E. Holcomb, Francesco Regazzoni 0001, Ulrich Rührmair, Patrick Schaumont |
CCS | 1 |
| 2019 | Analysis of Circuit Aging on Accuracy Degradation of Deep Neural Network AcceleratorabstractDeep neural networks have achieved phenomenal successes in vision recognition tasks, which motivate the deployment of deep learning in portable and smart wearable devices. To overcome the fundamental challenges of power and resource limitation, application-specific integrated circuit accelerators have emerged to compact the model and use lower precision arithmetic to increase the throughput of computation with reduced power consumption. Although very high energy efficiency has been achieved by removing redundant weights, compressing data and even sacrificing timing margin, such trend in hardware acceleration that pushes the deep learning systems to the error threshold can be disastrous for the tasks they performed due to failure or degraded performance of circuit components. Concerned by the lack of attention on the evolving unreliability effects in artificial intelligent accelerators implemented by the continuously scaled CMOS technology, this paper is the first to evaluate the effect of circuit aging on performance degradation of deep learning accelerator. Our findings indicate that DNN system running at their peak throughput rate can experience up to 84% accuracy drop after a year of aging and the accumulation of errors aggravates with the depth of learning. It is also found that relaxation of throughput rate can slow down the loss of classification accuracy considerably. Wenye Liu, Chip-Hong Chang |
ISCAS | 2 |
| 2019 | High-Speed True Random Number Generator Based on Differential Current Starved Ring Oscillators with Improved Thermal StabilityabstractTrue random number generator (TRNG) is a hardware security primitive that has been widely used in cryptography, Monte-Carlo simulation, video games, etc. with increasing significance in security solutions addressing emerging cyber-physical threats. This paper presents a new design of TRNG based on temperature insensitive frequency deviation extracted from a pair of matched current starved ring oscillators (CS-ROs). The proposed TRNG can maintain the entropy rate of the output bit stream over a broad range of temperatures. The CS-RO generates larger jitter noise than the regular ring oscillator, but consumes much less power for the same number of inverter stages. Monte Carlo simulated results based on a commercial 65 nm 1.2V CMOS process technology shows that it can output a random bit sequence at a rate of 485 Mbps with only 1.52 pJ of energy dissipation per bit. The proposed TRNG passes the NIST cryptographic randomness tests and autocorrelation test for the bit streams generated at temperature ranges from −40°C to 120°C. Yuan Cao 0003, Chip-Hong Chang, Xiaoli Ji |
ISCAS | 3 |
| 2019 | An In-Pixel Gain Amplifier Based Event-Driven Physical Unclonable Function for CMOS Dynamic Vision SensorsabstractIn this paper, a novel in-pixel event-driven physical unclonable function (PUF) is presented for the rapidly developed CMOS dynamic vision sensor (DVS). Different from traditional widely reported PUF implementations with additional dedicated silicon area, power consumption and peripheral circuitries, the proposed implementation extracts PUF based on the original gain amplifier existing in the mainstream DVS pixel, which is necessary to amplify the front-end logarithmic photoreceptor's relatively weak output signal, according to the ratio of the in-pixel capacitor pair. With any ON/OFF event generated and the corresponding DVS pixel fired asynchronously, the DVS pixel's own gain amplifier will be reset in order to capture the next possible event. Due to the inevitable variation of the semiconductor fabrication process, the reset voltages of different DVS pixels' gain amplifiers are slightly different. A bidirectional counter based analog-to-digital converter is customized to digitize the successively fired pixel pair with the sign bit representing the PUF bit (i.e. the reset voltages' difference). Moreover, the proposed implementation is validated using a standard 0.18μm CMOS process in Cadence. According to the extensive post-layout simulation results, the uniqueness is calculated to be 49.97%. With the operating temperature varying from -40°C to 120°C and supply voltage varying from 1.7V to 2.1V, the worst-case reliability is reported to be 96.48% and 97.27%, respectively. Meanwhile, its superior randomness is also verified using the NIST test suite. Biyin Wang, Xiaojin Zhao, Chip-Hong Chang |
ISCAS | 4 |
| 2019 | A Reliable Physical Unclonable Function Based on Differential Charging CapacitorsabstractPhysical Unclonable Function (PUF) is an emerging security primitive for cryptography applications. However, achieving a very high reliability against the environmental variations remains a main challenge in PUF design and a key barrier for its commercialization. This paper presents a new PUF design based on the charging of a symmetric MOS capacitor pair by constant current with cross-coupled positive feedback inverters. The proposed weak PUF features high raw response reliability against variations in power supply and temperature without power-up reset noise and other issues due to the power-down and up of an array of cells. Extensive Monte-Carlo simulations have been performed using a standard 110nm CMOS process technology. The simulated results show an almost ideal uniqueness of 50.03% and superior reliability of 97.70% over a temperature range from 0 °C to 80 °C, and 96.20% with the supply voltage varies from 1.2 V to 1.8 V. The response bit can be generated at a rate of 27.78 Mbps with an average power consumption of 20.86 μW at 1.5V, and the energy consumption is only 750 fJ/bit. Wei Guo 0018, Chip-Hong Chang, Yuan Cao 0003, Shaojun Wei, Shouyi Yin, Chenchen Deng, Leibo Liu, Fan Zhang 0044 |
ISCAS | 3 |
| 2019 | Emerging Attacks and Solutions for Secure Hardware in the Internet of ThingsabstractThe fourteen papers in this special section explore software solutions for secure hardware in the Internet of Things (IoT). It could well be argued that the emerging IoT, together with the two long-standing trends of pervasive and ubiquitous computing, constitutes one of the most massive civil endeavors in the history of mankind. While it promises outstandingly positive usability and convenience effects, its implications for security and privacy are less clear. The vision of billions of low-cost, lightweight, and highly interconnected endpoints certainly rises a host of pressing issues to both cryptographers and system designers. Ideally, these should be resolved prior to a large-scale deployment of the IoT, and before its underlying infrastructure and standards have been established. Chip-Hong Chang, Marten van Dijk, Ulrich Rührmair, Mark Tehranipoor |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2019 | Reliable and Modeling Attack Resistant Authentication of Arbiter PUF in FPGA Implementation With Trinary Quadruple ResponseabstractField programmable gate array (FPGA) is a potential hotbed for malicious and counterfeit hardware infiltration. Arbiter-based physical unclonable function (A-PUF) has been widely regarded as a suitable lightweight security primitive for FPGA bitstream encryption and device authentication. Unfortunately, the metastability of flip-flop gives rise to poor A-PUF reliability in FPGA implementation. Its linear additive path delays are also vulnerable to modeling attacks. Most reliability enhancement techniques tend to increase the response predictability and ease machine learning attacks. This paper presents a robust device authentication method based on the FPGA implementation of a reliability enhanced A-PUF with trinary digit (trit) quadruple responses. A two flip-flop arbiter is used to produce a trit for metastability detection. By considering the ordered responses to all four combinations of first and last challenge bits, each quadruple response can be compressed into a quadbit that represents one of the five classes of trit quadruple response with greater reproducibility. This challenge-response quadruple classification not only greatly reduces the burden of error correction at the device but also enables a precise A-PUF model to be built at the server without having to store the complete challenge-response pair (CRP) set for authentication. Besides, the real challenge to the A-PUF is generated internally by a lossy, nonlinear, and irreversible maximum length signature generator at both the server and device sides to prevent the naked CRP from being machine learned by the attacker. The A-PUF with short repetition code of length five has been tested to achieve a reliability of 1.0 over the full operating temperature range of the target FPGA board with lower hardware resource utilization than other modeling attack resilient strong PUFs. The proposed authentication protocol has also been experimentally evaluated to be practically secure against various machine learning attacks including evolutionary strategy covariance matrix adaptation. Siarhei S. Zalivaka, Alexander A. Ivaniuk, Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2018 | ASHES 2018- Workshop on Attacks and Solutions in Hardware SecurityabstractAs in the successful first edition, the second Workshop on Attacks and Solutions in Hardware Security (ASHES) 2018 deals with all aspects of hardware security. Among others, this year, the workshop particularly highlights emerging techniques and methods as well as recent application areas within the field. These include new attack vectors, attack countermeasures, and novel designs and implementations on the methodological side, as well as the Internet of Things, automotive security, smart homes, pervasive and wearable computing on the applications side. In order to meet the requirements of these rapidly developing subareas, ASHES calls for paper submissions in four categories: 1) classical full papers; 2) classical short papers; 3) systematization of knowledge papers which overview, structure, and categorize a subarea; and 4) wild and crazy papers whose purpose is rapid dissemination of promising, potentially game-changing ideas. Chip-Hong Chang, Jorge Guajardo, Daniel E. Holcomb, Francesco Regazzoni 0001, Ulrich Rührmair |
CCS | 1 |
| 2018 | A Fully Digital Physical Unclonable Function Based Temperature Sensor for Secure Remote SensingabstractTurnkey solutions that combine energy-efficient remote sensing and secure communication of telemetry are desirable in data collection, risk control and situation appraisal with the large scale deployment of resource constrained Internet of Things devices. In this paper, a new low-cost physical unclonable function (PUF) based temperature sensor for secure remote temperature sensing is proposed. The design exploits the approximately linear positive temperature coefficient of CMOS inverter in super-threshold operation to calibrate the running frequency of ring oscillator (RO) in a reconfigurable RO PUF at different temperature. The RO frequency corresponding to the sensed temperature is fed into a randomizer seeded by the input challenge to select new RO pairs for comparison to generate a random, unique and physically unclonable digital tag, which is valid for a selected input challenge to a target device at a particular temperature. Using only standard logic cells and a very simple structure, the proposed temperature sensor can be easily implemented on FPGA and integrated into other digital systems. It protects the integrity of the sensed information by preventing falsified sensor data and masquerade sensing node. The FPGA implementation of our proposed design has demonstrated the feasibility of making a trust temperature telemetry system out of PUF. Yuan Cao 0003, Yunyi Guo, Benyu Liu, Min Zhu 0001, Chip-Hong Chang |
ICCCN | 6 |
| 2018 | A Sub-pico Joules Per Bit Robust Physical Unclonable Function Based on Subthreshold Voltage ReferencesabstractLow power, lightweight and robust physical unclonable function (PUF) is a sought-after for IoT device identification/authentication. This paper presents a low power PUF design with high reliability against temperature and supply voltage variations. A response bit is extracted by comparing a pair of identically designed subthreshold voltage references. The voltage difference due to device mismatch is digitized and registered in a bidirectional counter, which can be used to identify and filter out the unstable response bits. The readout circuit works in tandem with the proposed double sampling technique to reduce the bias of components that are not the main entropy source of response bits. The proposed design is evaluated by extensive simulation using standard 65 nm CMOS process. It consumes merely 0.16 pJ/bit. The simulated uniqueness is an almost ideal 50.03%. Due to the intrinsic stability of voltage references, the reliability of its native response is 98.17% for the supply voltage variation from 1 V to 1.4 V and 97.60% for the temperature variation from 0 ° C to 80 °C. The generated response bitstream has passed both the autocorrelation test and NIST randomness test. Yuan Cao 0003, Chip-Hong Chang, Wenhan Zheng, Xiaojin Zhao |
ISCAS | 2 |
| 2018 | Active IC Metering of Digital Signal Processing Subsystem with Two-Tier Activation for Secure Split TestabstractActive integrated circuit (IC) metering is a class of hardware security protocols that enables the designer to track the number of chips produced from the same mask and remotely activate only the desired ones. This paper reviews existing IC metering approaches to incorporate the advantages of individual methods into a secure functional lock on digital signal processing submodule of wireless communication system to avoid legitimate channel exploitation and the risk of deploying unreliable out-of-specs gray market ICs. Our method makes use of aging-sensitive physical unclonable function to enable a two-tier activation of ICs in split test flow to track chip supply after production tests. Extraneous states are inserted into the state-space mapping of digital signal processing submodule as opposed to controller to provide a stronger state dependency on datapath and input signal. The scheme is illustrated experimentally on a pulse shaping filter of the transmitter for a wireless communication system. Sumedh Dhabu, Wenye Liu, Chip-Hong Chang |
ISCAS | 4 |
| 2018 | Securing IoT Monitoring Device using PUF and Physical Layer AuthenticationabstractIoT is rapidly becoming a reality. Forecasts predict more than 20 billion connected devices in 2020. These devices bring many benefits, but securing them in IoT environment can be a quandary. With the advent of technology, it is very easy for an adversary to clone a device and replace it, or tamper the data. In the context of wireless communications in IoT, the definition of message authentication should be extended to include verification of the device along with the integrity of the message it produced. In this paper we propose a device- and data-dependent physical layer authentication scheme by using a device-specific, dynamically variable key to generate a data-dependent tag. This tag is embedded in the data transmission using an information hiding scheme to reliably extract it at the receiver, and without compromising the performance of the underlying wireless communication system. Simulation results show that our scheme can achieve high authentication rate while rejecting the tampered transmissions in typical noisy communication channel. Sumedh Dhabu, Chip-Hong Chang |
ISCAS | 3 |
| 2018 | New Hardware and Power Efficient Sporadic Logarithmic Shifters for DSP ApplicationsabstractShifting an input data by variable amounts is commonly found in arithmetic operations, data encoding and bit-indexing. Although some shift amounts along the entire shift range are not required, the typical realization by a full-range logarithm shifter requires full implementation and therefore suffers from complexity and power overhead. In this paper, the notion of sporadic logarithmic shifter (SLS) is introduced for the first time, and a new design methodology is proposed for its optimization. By reusing parts of existing substructure of conventional logarithmic shifter or post-multiplexing the hardwired shifts, contagious subranges of desirable shift amounts are successively realized. Synthesis results on 8-bit and 16-bit SLSs show average application-specified integrated circuit area and power savings of up to 73.24% and 63.90%, respectively, over conventional logarithmic shifters. In addition, by applying the proposed SLSs to the discrete cosine transform (DCT) architecture, at least 47.5% area savings and 2.9% power savings can be achieved over two constant multipliers-based DCT architectures reported in the literature. Jiajia Chen 0002, Chip-Hong Chang, Juan Zhao 0005, Susanto Rahardja |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | A New Accurate and Fast Homography Computation Algorithm for Sports and Traffic Video AnalysisabstractHomography has wide applications in aerial photographic surveys, camera calibration, traffic scene analysis, and sports science, such as player and team performance evaluation. Unlike the mainstream homography that utilizes points as matching features, homography estimation for sports and traffic video can achieve higher accuracy and speed by utilizing straight lines in the scenes, which convey more information than points. Owing to the more stringent requirement of accuracy and computational speed for advanced video analysis, this paper presents a novel homography computational algorithm. Three major novelties are proposed and validated, which are multiple points Hough transform for straight line extraction, correspondence initialization by angle to estimate a set of quasi-optimal solutions, and the feature correspondences optimization to achieve a minimized error using genetic algorithm. With these contributions, the experiments have shown that the proposed algorithm can improve the homography computational accuracy by up to 130% and reduce the processing time by up to 96% over the state-of-the-art algorithms for the same purposes. ShuMin Liu, Jiajia Chen 0002, Chip-Hong Chang, Ye Ai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | ASHES 2017: Workshop on Attacks and Solutions in Hardware SecurityabstractThe workshop on "attacks and solutions in hardware security" (ASHES) deals with all aspects of hardware security, including any recent attacks and solutions in the area. Besides mainstream research in hardware security, it also covers new, alternative or emerging application scenarios, such as the internet of things, nuclear weapons inspections, satellite security, or consumer and supply chain security. It also puts some focus on special purpose hardware and novel methodological solutions, such as particularly lightweight, small, low-cost, and energy-efficient devices, or even non-electronic security systems. Finally, ASHES welcomes any theoretical works that systematize and structure the area, and so-called "Wild-and-Crazy" papers that describe and distribute seminal ideas at an early conceptual stage to the community. Chip-Hong Chang, Marten van Dijk, Farinaz Koushanfar, Ulrich Rührmair, Mark Tehranipoor |
CCS | 1 |
| 2017 | A new watermarking scheme on scan chain ordering for hard IP protectionabstractConstraint-based watermarking has been proven as an effective means for hardware intellectual property (IP) protection. Scan chain further facilitates field authentication of the watermarks embedded in densely integrated IP cores and serves as an added layer of tamper-evident protection if it is also watermarked. In this paper, we propose a new scan-chain based watermarking scheme that can be used standalone as a reasonably robust copyright protection of hard IP core. During the process of scan chain ordering for test power optimization, the watermark bits constrain the cost of one connection style over the other for certain pairings of scan cells in the minimization process depending on the output of the IP core under the chosen test vector for watermark verification. Thus, the watermark is better obfuscated since both connection styles are possible to be selected at the watermarked locations. Typical attack scenarios are discussed and our experimental results show that strong authorship can be achieved with low overhead on test power incurred. Xiaonan Huang, Aijiao Cui, Chip-Hong Chang |
ISCAS | 3 |
| 2017 | A new write-contention based dual-port SRAM PUF with multiple response bits per cellabstractUsing power-up reset to reuse Static Random Access Memory (SRAM) as physical unclonable function (PUF) suffers from two major constraints. First, it has limited entropy of only one response bit per cell. Second, as power-up reset has a global effect, extra storage is needed to temporarily buffer the original SRAM content before switching to PUF mode. This process has introduced extra area-power overhead as well as potential security leakage. The emerging dual-port (DP) SRAM cell offers attractive multiple access capability for low-power and high-speed memory transfer. By levering the forbidden contention state in one of its four multiple access modes, a new DP-SRAM based PUF is proposed to generate two independent response bits per cell and limit data buffering to only those cell content addressed by the challenge. Simulation results based on PTM 22nm LP CMOS model has corroborated its superior reliability, uniqueness and randomness over and above these meritorious advantages. Chao Qun Liu, Chip-Hong Chang |
ISCAS | 3 |
| 2017 | 20 Years of research on intellectual property protectionabstractVLSI intellectual property (IP) reuse based design methodology was adopted by the semiconductor industry in the early 1990's and how to protect design IPs from piracy and misuse has since been a challenging problem. 2017 marks the 20th anniversary of the IP protection development and working group was founded and the first series of IP watermarking papers were published. In this paper, we survey the efforts from industry, government, and academia on securing the design IPs in the past 20 years with focus on development from academia side. Miodrag Potkonjak, Gang Qu 0001, Farinaz Koushanfar, Chip-Hong Chang |
ISCAS | 4 |
| 2017 | Current mirror array: A novel lightweight strong PUF topology with enhanced reliabilityabstractIn this work, we present a novel design of physical unclonable function (PUF) based on the topology of current mirror array (CMA). The proposed strong PUF exploits the randomness in the current mirror transistors and generates the response bit by comparing the accumulated current values. Thresholding and reference current are used to increase the reliability of the proposed PUF without requiring additional hardware resources. The proposed PUF structure is also analysed in terms of its difficulty of model building for measurement-prediction attack. Measurement results on 0.35μm test chips demonstrate that the proposed PUF outperforms other state-of-the-art designs with smaller area/bit of 9 × 10−36μm2 and lower native bit error rate (BER) of 0.16%. Zheng Wang 0027, Yi Chen 0012, Aakash Patil, Chip-Hong Chang, Arindam Basu |
ISCAS | 4 |
| 2017 | Low-cost fortification of arbiter PUF against modeling attackabstractArbiter Physical unclonable function (A-PUF) with exponential number of challenges is an ideal candidate to realize lightweight and robust device authentication in Internet of Things applications. Unfortunately, it is particularly difficult to attain highly reliable responses and increase its modeling attack resistance simultaneously. This paper presents an approach to reduce the vulnerability of A-PUF to machine learning attacks without compromising its high reliability and uniqueness. It utilizes a multiple input signature register (MISR) to process the input challenges. Our experiment results show that the accuracy of predicting the responses of a MISR augmented 128-stage arbiter PUF in FPGA implementation by support vector machine and gradient boosting learning algorithms with a training set of 100,000 challenge-response pairs has reduced drastically from 98% to 50%. If design-for-testability is mandatory, the MISR can be reconfigured from an existing built-in logic block observer, making this approach virtually free. Otherwise, the MISR carries a negligible hardware overhead of only 0.4% of the total available resources in an Xilinx ZC706 FPGA chip. Siarhei S. Zalivaka, Alexander A. Ivaniuk, Chip-Hong Chang |
ISCAS | 3 |
| 2017 | Static and Dynamic Obfuscations of Scan Data Against Scan-Based Side-Channel AttacksabstractDue to the fallibility of advanced integrated circuit (IC) fabrication processes, scan test has been widely used by cryptographic ICs to provide high fault coverage. Full controllability and observability offered by the scan design also open out the trapdoor to side-channel attacks. To better resist signature attacks on scan testable cryptochip, we propose to fortify the key and lock method by the static obfuscation of scan data. Instead of spatially reshuffling the scan cells, the working mode of some scan cells is altered to jumble up the scan data when the scan test is performed with an incorrect test key. However, when the plaintext is fed directly through the primary inputs for test efficiency, the static obfuscation of scan data is inadequate as demonstrated by a new test-mode-only signature attack (TMOSA) proposed in this paper. To thwart TMOSA, a new countermeasure based on the dynamic obfuscation of scan data is proposed. By cyclically shifting the incorrect test key throughout the test phase, the blocking cells due to the mismatched bits of the test key are made to move temporally to dynamically obfuscate the scan data. This latter scheme is unconditionally resilient against TMOSA and all other known scan-based attacks while preserving the merits of high testability and low area overhead compared with other countermeasures. Aijiao Cui, Yanhui Luo, Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2017 | A Scaling-Assisted Signed Integer Comparator for the Balanced Five-Moduli Set RNS 2n-1, 2n, 2n+ 1, 2n+1-1, 2n-1-1abstractSigned integer comparison occurs prevalently in process control, sorting, address decoding, and conditional branching. Existing algorithms for magnitude comparison in RNS are based either on parity check or partial reverse conversion. With separately designed RNS sign detectors, they can also be used to compare residue representations of signed integers. In this paper, a radically different approach to this problem is proposed for the five-moduli set {2n- 1, 2n, 2n+1, 2n+1- 1, 2n-1- 1}. The signs of the operands in comparison, as well as their difference are detected after scaling by a factor of (22n-1)(2n-1-1). The resulting finite series in the composite modulus channel is further factored into parallel carry-saved additions in the existing mod 2n and mod 2n+1- 1 modulus channels, thus reducing the sizes of modulo adders from 5n bits to n and n+1 bits. Upon detecting the signs of the operands and their difference, the relation is inference with a small fraction of logic gates. Our synthesis results in 65-nm CMOS technology show that the proposed design is 36.9% smaller, 7.6% faster, and 45.5% more energy efficient than the best Chinese remainder theorem-II-based magnitude comparator and at least 12.9% smaller, 7.3% faster, and 20.8% more energy efficient than the best reverse-conversion-based implementation of signed integer comparator for the same five-moduli set. Sachin Kumar 0006, Chip-Hong Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Multi-valued Arbiters for quality enhancement of PUF responses on FPGA implementationabstractOne main problem encountered in the FPGA implementation of Arbiter based Physical Unclonable Function (A-PUF) is the response instability caused by the metastability of delay flip-flop. This paper presents a new multi-arbiter approach to extract more entropy to extend the number of response bits to a single challenge. New multi-arbiter schemes based on the insertion of either a four-flip-flop arbiter or SR latch arbiter after each pair of multiplexers in the configurable paths are proposed to detect the metastable state when two copies of test pulse arrive at the arbiter inputs almost simultaneously. The detected metastable states are distinguishable by the encoded multiple valued outputs of the arbiter. The codes corresponding to the metastable states collectively form a deterministic ternary state that can be recoded to one of the stable states to improve the uniqueness and reliability of the PUF. Our analysis shows that the proposed design can generate robust and reliable challenge-response pairs with a uniqueness of 0.4982 and a reliability of 0.9985 at the expense of a relatively small FPGA resource overhead. Siarhei S. Zalivaka, Alexander V. Puchkov, Vladimir P. Klybik, Alexander A. Ivaniuk, Chip-Hong Chang |
ASP-DAC | 5 |
| 2016 | A VLSI-efficient signed magnitude comparator for { 2n-1, 2n, 2n +2n+1-1} RNSabstractComparison of residue representations in signed Residue Number System (RNS) involves sign detection and magnitude comparison. Both are difficult operations in RNS. This paper proposes a new signed magnitude comparator for the three-moduli set RNS {2n-1, 2n, 2n+1-1}. Two subrange identifiers are computed to simplify sign detection and accelerate the magnitude comparison without requiring full reverse conversion, large modulo adders or lookup tables. Synthesis results in 65 nm CMOS standard cell implementation show that it even outperforms the most efficient unsigned magnitude comparator for equally balanced special three-moduli set by significant margin in terms of area, delay and power consumption. Sachin Kumar 0006, Chip-Hong Chang |
ISCAS | 2 |
| 2016 | A low-voltage, low power STDP synapse implementation using domain-wall magnets for spiking neural networksabstractOnline, real-time learning in neuromorphic circuits have been implemented through variants of Spike Time Dependent Plasticity (STDP). Current implementations have used either floating-gate devices or memristors to implement such learning synapses together with non-volatile storage. However, these approaches require high voltages (≈ 3-12V) for weight update and entail high energy for learning (≈ 4-30pJ/write). We present a domain wall memory based low-voltage, low-energy STDP synapse that can operate with a power supply as low as 0.8V and update the weight at ≈ 40fJ/write. Device level simulations are performed to prove its feasibility. Its use in associative learning is also demonstrated by using neurons with dendritic branches to classify spike patterns from MNIST dataset. Govind Narasimman, Subhrajit Roy, Xuanyao Fong, Kaushik Roy 0001, Chip-Hong Chang, Arindam Basu |
ISCAS | 5 |
| 2016 | A Non-Iterative Multiple Residue Digit Error Detection and Correction Algorithm in RRNSabstractError detection and correction code based on Redundant Residue Number System (RRNS) has a unique advantage that arithmetical processing errors can also be corrected. Existing algorithms for multiple residue digit error correction in RRNS require either large modulo operations to decode the magnitude of the residue digits or variable number of iterative computations to identify the locations of erroneous residue digits. This paper presents an efficient syndrome based multiple residue digit error detection and correction algorithm in RRNS. The received residue digits are divided into three groups from which seven error location categories are defined for all combinations of residue digit errors of any legitimate moduli set. Their error magnitudes are abstracted into three syndromes, which are used to identify the exact error location category and retrieve the residue digit errors concurrently from six lookup tables in no more than three cycles. The syndromes can be computed in parallel by small modulo subtractors, and their uniqueness criteria are proved to be easily fulfilled. The redundant residue space due to the uniqueness criteria can also be relaxed by increasing the number of syndrome computations and lookup tables used for the error decoding according to the number of information moduli. Thian Fatt Tay, Chip-Hong Chang |
IEEE Trans. Computers | 2 |
| 2016 | A New Paradigm of Common Subexpression Elimination by Unification of Addition and SubtractionabstractThis paper makes a paradigm shift in the assumed notion of common subexpressions for complexity reduction of multiple constant multiplications implementation. Our proposed unified adder/subtractor (UAS)-based common subexpression elimination (CSE) algorithm is inspired by the recent advancement in complex arithmetic component mapping for datapath synthesis of digital systems. A dedicated UAS operator is designed at gate level to achieve arithmetic reduction for concurrent computation of the sum and difference of two input signals. To maximize computation reuse, dual subexpression is defined to enable a UAS to be shared by the otherwise incompatible odd and even common subexpressions. The three different types of common subexpression are uniquely encoded by a quadruple in the proposed data structure. Constant coefficients are represented by signed digits in Cartesian coordinate system from which nonoverlapping pairs of nonzero digits are parsed for dual, even, and odd subexpressions to maximize the reuse of all three types of arithmetic resources. The effectiveness of our proposed UAS-based CSE in overcoming the complexity reduction bottleneck are demonstrated by comparing the synthesis results obtained from six benchmark finite impulse response filters, an electroencephalogram filter bank, fast Fourier transform, and discrete cosine transform multipliers designed by ten algorithms. The results show a noteworthy 27.2% reduction in area-time complexity of our method over the baseline canonical signed digit implementation. Our solutions are also more power efficient, with average power saving of 12.0% over those designed by other algorithms in comparison. Jiatao Ding, Jiajia Chen 0002, Chip-Hong Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | DW-AES: A Domain-Wall Nanowire-Based AES for High Throughput and Energy-Efficient Data Encryption in Non-Volatile MemoryabstractBig-data storage poses significant challenges to anonymization of sensitive information against data sniffing. Not only will the encryption bandwidth be limited by the I/O traffic, the transfer of data between the processor and the memory will also expose the input-output mapping of intermediate computations on I/O channels that are susceptible to semi-invasive and non-invasive attacks. Limited by the simplistic cell-level logic, existing logic-in-memory computing architectures are incapable of performing the complete encryption process within the memory at reasonable throughput and energy efficiency. In this paper, a block-level in-memory architecture for advanced encryption standard (AES) is proposed. The proposed technique, called DW-AES, maps all AES operations directly to the domain-wall nanowires. The entire encryption process can be completed within a homogeneous, high-density, and standby-power-free non-volatile spintronic-based memory array without exposing the intermediate results to external I/O interface. Domain-wall nanowire-based pipelining and multi-issue pipelining methods are also proposed to increase the throughput of the baseline DW-AES with an insignificant area overhead and negligible difference on leakage power and energy consumption. The experimental results show that DW-AES can reduce the leakage power and area by the orders of magnitude compared with existing CMOS ASIC accelerators. It has an energy efficiency of 22 pJ/b, which is 5× and 3× better than the CMOS ASIC and memristive CMOL-based implementations, respectively. Under the same area budget, the proposed DW-AES achieves 4.6× higher throughput than the latest CMOS ASIC AES with similar power consumption. The throughput improvement increases to 11× for pipelined DW-AES at the expense of doubling the power consumption. Yuhao Wang 0002, Leibin Ni, Chip-Hong Chang, Hao Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | A New Fast and Area-Efficient Adder-Based Sign Detector for RNS {2n-1, 2n, 2n+1}abstractThe moduli set {2n- 1, 2n, 2n+ 1} has been widely used in residue number system (RNS)-based computations. Its sign extraction problem, albeit fundamentally important in magnitude comparison and other difficult algorithms in RNS, has received considerably less attention than its scaling and reverse conversion problems. This brief presents a new algorithm for the design of a fast adder-based sign detector. The circuit is greatly simplified by shrinking the dynamic range to eliminate large modulo operations with the help of the new Chinese remainder theorem. Our synthesis results with the 65-nm CMOS standard cell library show that the proposed design outperforms all the existing adder-based sign detectors reported for this moduli set in area and speed for n ranges from 5 to 25 in the step of 5. Sachin Kumar 0006, Chip-Hong Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Erratum to "Efficient VLSI Implementation of 2n Scaling of Signed Integer in RNS {2n-1, 2n, 2n+1}"abstractThe authors would like to point out the following correction in the total power consumptions recorded in Table IV of the published article[1]. The total power consumptions of the proposed design and the design[2]in comparison should be corrected as shown inTable I. Thian Fatt Tay, Chip-Hong Chang, Jeremy Yung Shern Low |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A new unified modular adder/subtractor for arbitrary moduliabstractEfficient modular adders and subtractors for arbitrary moduli are key booster of computational speed for high-cardinality Residue Number Systems as they rely on arbitrary moduli set to expand the dynamic range. This paper proposes a new unified modular adder/subtractor that possesses a regular structure for any modulus. Compared to the latest modular adder/subtractor, which works for modulus in the forms of 2n±k, the proposed design is on average 10.81% faster and consumes 15.85% less hardware area and 2.51% lower power for n ranging from 4 to 8. Thian Fatt Tay, Chip-Hong Chang |
ISCAS | 2 |
| 2015 | Public key protocol for usage-based licensing of FPGA IP coresabstractApplication developers are now turning to field-programmable gate array (FPGA) devices for solutions of small to medium volume due to its post-fabrication flexibility. Unfortunately, the existing upfront intellectual property (IP) licensing model for FPGA based third-party IP cores is economically unattractive. The IP bitstreams in transaction are also vulnerable to cloning, misappropriation and reverse engineering. This paper proposes a secure pay-per-use licensing protocol to avoid complicated communication flow and high implementation cost, while preventing the IP rights from being compromised or abused by all parties involved. The protocol guarantees the confidentiality and integrity of the security-critical components and forbids the implementation of licensed IP cores on gray market or counterfeit chips. The public-crypto based core installation module used to self-configure the licensed IP cores occupies only limited FPGA fabrics temporarily. Chip-Hong Chang |
ISCAS | 2 |
| 2015 | Statistical analysis and design of 6T SRAM cell for physical unclonable function with dual application modesabstractApart from performance and power efficiency, security is another critical concern in the modern memory sub-system design. SRAM, which is routinely used as a data preservation component, has now been developed into an effective primitive known as Physical Unclonable Function (PUF) for cryptographic key generation to protect the sensitive local information. Considering the constraints of hardware resource on embedded systems, it is desirable to have an SRAM used both as a regular memory and a PUF to save the overheads of having these two functions implemented independently. Unfortunately, while process variations are the entropy sources for secure key generation, it impacts failure rates in memory-mode operations. This paper presents a statistical analysis on SRAM and provides an insight into how the SRAM cell geometry can be optimized to qualify it for both modes of operation simultaneously. Le Zhang 0001, Chip-Hong Chang, Zhi-Hui Kong, Chao Qun Liu |
ISCAS | 2 |
| 2015 | Design of Optimal Scan Tree Based on Compact Test Patterns for Test Time ReductionabstractScan tree architecture has been proposed to reduce the test application time of full scan chain by placing multiple scan cells in parallel. Most existing techniques rely on non-compact test pattern sets to construct the scan tree. However, they produce inefficient scan tree when highly compact test sets with few don't cares are used. In this paper, the depth of the scan tree based on approximate compatibility relation for completely specified test data set is analyzed probabilistically by modeling its construction as a vertex coloring problem. The upper bound of edges-per-vertex is computed and demonstrated to be a prime factor that limits the efficiency of scan tree construction based on both compatible and approximately compatible test data between two flip-flops. Inverse compatibility and aggressive approximate compatibility are then proposed to increase the edges-per-vertex for vertex coloring. The Q'-SD connection between two adjacent scan cells is exploited to implement the inverse compatibility with no cost or timing impact. To maintain the fault coverage, the missing faults under the tree scan mode can be detected by switching the same base architecture into the linear scan mode with negligible hardware overhead as shown by the experimental results on ISCAS89, ISCAS99 and LGSynth93 benchmark circuits. On average, the scan tree generated by our method reduces the test time of the full scan chain by 56.65 percent, and that of the scan tree designed by the approximate compatibility method by 39.18 percent under the same compact test sets. Linfeng Chen, Aijiao Cui, Chip-Hong Chang |
IEEE Trans. Computers | 3 |
| 2015 | A Low-Power Hybrid RO PUF With Improved Thermal Stability for Lightweight ApplicationsabstractRing oscillator (RO)-based physical unclonable function (PUF) is resilient against noise impacts, but its response is susceptible to temperature variations. This paper presents a low-power and small footprint hybrid RO PUF with a very high temperature stability, which makes it an ideal candidate for lightweight applications. The negative temperature coefficient of the low-power subthreshold operation of current starved inverters is exploited to mitigate the variations of differential RO frequencies with temperature. The new architecture uses conspicuously simplified circuitries to generate and compare a large number of pairs of RO frequencies. The proposed nine-stage hybrid RO PUF was fabricated using global foundry 65-nm CMOS technology. The PUF occupies only 250 μm2of chip area and consumes only 32.3 μW per challenge response pair at 1.2 V and 230 MHz. The measured average and worst-case reliability of its responses are 99.84% and 97.28%, respectively, over a wide range of temperature from -40 to 120 °C. Yuan Cao 0003, Le Zhang 0001, Chip-Hong Chang, Shoushun Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | Optimizating Emerging Nonvolatile Memories for Dual-Mode Applications: Data Storage and Key GeneratorabstractMemory-based physical unclonable functions (PUFs) have been studied and developed as powerful primitives to generate device-specific random keys, which can be used for various security applications. However, the existing memory-based PUFs need to safely buffer the data bits in the memory before it is used to produce random bits, resulting in additional area/energy consumption and potential data security issues. In this paper, we propose a new memory-based PUF that exploits the nonvolatility and random variability of emerging memory technologies to produce random bits. Unlike conventional implementations, the random bit generation process of our proposed PUF does not disturb the data bits already stored in the memory. To satisfy the quality requirements for both memory and PUF applications, we also propose a general method to find the optimal design point of emerging nonvolatile memory (eNVM)-based PUF. An illustrative design using spin-transfer torque magnetic RAM exhibits desirable results using our method. Compared to the conventional types of memory-based PUFs, eNVM-based PUFs features enhanced security as cryptographic primitives and lower area and energy cost as data storage. Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | Highly Reliable Spin-Transfer Torque Magnetic RAM-Based Physical Unclonable Function With Multi-Response-Bits Per CellabstractMemory-based physical unclonable function (MemPUF) has gained tremendous popularity in the recent years to securely preserve secret information in computing systems. Most MemPUFs in the literature have unreliable bit generation and/or are incapable of generating more than one response-bit per cell. Hence, we propose a novel MemPUF exploiting the unique characteristics of spin-transfer torque magnetic RAM (STT-MRAM) that can overcome these issues. Bit generation in our STT-MRAM-based MemPUF is stabilized using a novel automatic write-back technique. In addition, the alterability of the magnetic tunneling junction state is exploited to expand the response-bit capacity per cell. Our analysis demonstrated the advantage of our scheme in reliability enhancement (bit-error rate from ~10-1to ~10-6in the worst case under varying conditions) and response-bit capacity per cell improvement (from 1 to 1.48 bit). In comparison with the conventional MemPUFs, our approach is also better in terms of the average chip area and energy for producing a response-bit. Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Area-efficient and fast sign detection for four-moduli set RNS {2n -1, 2n, 2n +1, 22n +1}abstractSign detection is a necessary but non-trivial operation in Residue Number System (RNS) for many digital signal processing applications. Efficient sign detector for the three-moduli set RNS {2n-1,2n, 2n+1} has been proposed, but the problem remains unsolved for its extended four moduli sets. This paper presents a new sign detection algorithm dedicated to {2n-1,2n, 2n+1,22n+1} RNS that has a wider dynamic range and higher parallelism. Our approach exploits the number theoretic and multiplicative inverse properties in two-residue Chinese Remainder Theorem (CRT) and the New CRT II to halve the bit width of the modulo additions required by a complete reverse conversion. Our synthesis results show greater than 60% area reduction and more than 40% speedup for n = 2 to 5 compared with using its most efficient reverse converter for sign detection. Chip-Hong Chang, Sachin Kumar 0006 |
ISCAS | 1 |
| 2014 | Design of programmable FIR filters using Canonical Double Based Number RepresentationabstractScalability of current programmable FIR filter design methods are severely limited by the huge search space for common subexpressions and the density of unique subexpressions over the complete range of integers of desirable precision. This paper presents the first attempt to solve this problem by means of Canonical Double-Based Number Representation (CDBNR). We address the representation sparsity of generic filter coefficients by developing a simplified CDBNR search algorithm. The statistics generated for all double base products of a given coefficient word length are used to maximize the sharing of arithmetic operators and reduce the multiplexing cost. The effectiveness and scalability of the proposed design algorithm are demonstrated using two design examples. For the 8-bit programmable filter example implementable by two latest design methods, our proposed solution saves about 24% of arithmetic operator and multiplexer costs for large filter. Jiajia Chen 0002, Chip-Hong Chang |
ISCAS | 2 |
| 2014 | A new algorithm for single residue digit error correction in Redundant Residue Number SystemabstractThis paper presents a new algorithm for the correction of single residue digit error in Redundant Residue Number System. The location and magnitude of error can be extracted directly from a minimum size lookup table. This is made possible by the introduction of a new syndrome, which is proven to be unique for every different residue digit error with two criteria imposed on the choice of redundant moduli. The erroneous residue digit can be corrected by deducting the error digit retrieved from the lookup table indexed by the syndrome. Our proposed algorithm compares favorably against existing single residue digit correction algorithms, and is more amenable to hardware implementation. Thian Fatt Tay, Chip-Hong Chang |
ISCAS | 2 |
| 2014 | Highly reliable memory-based Physical Unclonable Function using Spin-Transfer Torque MRAMabstractIn recent years, Physical Unclonable Function (PUF) based on the inimitable and unpredictable disorder of physical devices has emerged to address security issues related to cryptographic key generation. In this paper, a novel memory-based PUF based on Spin-Transfer Torque (STT) Magnetic RAM, named as STT-PUF, is proposed as a key generation primitive for embedded computing systems. By comparing the resistances of STT-MRAM memory cells which are initialized to the same state, response bits can be generated by exploiting the inherent random mismatches between them. To enhance the robustness of response bits regeneration, an Automatic Write-Back (AWB) technique is proposed without compromising the resilience of STT-PUF against possible attacks. Simulations show that the proposed STT-PUF is able to produce raw response bits with uniqueness of 50.1% and entropy of 0.985 bit per cell. The worst-case Bit-Error Rate (BER) under varying operating conditions is 6.6 × 10-6. Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001 |
ISCAS | 3 |
| 2014 | A 0.7 V low-power fully programmable Gaussian function generator for brain-inspired Gaussian correlation associative memory
Chip-Hong Chang, Arindam Basu, Liter Siek |
Neurocomputing | 2 |
| 2014 | A Blind Dynamic Fingerprinting Technique for Sequential Circuit Intellectual Property ProtectionabstractDesign fingerprinting is a means to trace the illegally redistributed intellectual property (IP) by creating a unique IP instance with a different signature for each user. Existing fingerprinting techniques for hardware IP protection focus on lowering the design effort to create a large number of different IP instances without paying much attention on the ease of fingerprint detection upon IP integration. This paper presents the first dynamic fingerprinting technique on sequential circuit IPs to enable both the owner and legal buyers of an IP embedded in a chip to be readily identified in the field. The proposed fingerprint is an oblivious ownership watermark independently endorsed by each user through a blind signature protocol. Thus, the authorship can also be proved through the detection of different user's fingerprints without the need to separately embed an identical IP owner's signature in all fingerprinted instances. The proposed technique is applicable to both application-specific integrated circuit and field-programmable gate array IPs. Our analyses show that the fingerprint is immune to collusion attack and can withstand all perceivable attacks, with a lower probability of removal than state-of-the-art FSM watermarking schemes. The probability of coincidence of a 32-bit fingerprint is in the order of 10-10and up to 103532-bit fingerprinted instances can be generated for a small design of 100 flip-flops. Chip-Hong Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | A Cluster-Based Distributed Active Current Sensing Circuit for Hardware Trojan DetectionabstractThe globalization of integrated circuits (ICs) design and fabrication has given rise to severe concerns on the devastating impact of subverted chip supply. Hardware Trojan (HT) is among the most dangerous threats to defend. The dormant circuit inserted stealthily into the chip by the advisory could steal the confidential information or paralyze the system connected to the subverted chip upon the HT activation. This paper presents a transient power supply current sensor to facilitate the screening of an IC for HT infection. Based on the power gating scheme, it converts the current activity on local power grid into a timing pulse from which the timing and power-related side channel signals can be externally monitored by the existing scan test architecture. Its current comparator threshold can be calibrated against the quiescent current noise floor to reduce the impacts of process variations. Postlayout statistical simulations of process variations are performed on the ISCAS'85 benchmark circuits to demonstrate the effectiveness of the proposed technique for the detection of delay-invariant and rarely switched HTs. Compared with the detection error rate of a 4-bit counter-based HT reported by an existing HT detection method using the path delay fingerprint, our method shows an order of magnitude improvement in the detection accuracy. Yuan Cao 0003, Chip-Hong Chang, Shoushun Chen |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | A Pragmatic Per-Device Licensing Scheme for Hardware IP Cores on SRAM-Based FPGAsabstractThe modus operandi of the upfront intellectual property (IP) licensing model is impractical for the developers of small to medium volume of field-programmable gate array (FPGA)-based applications. The FPGA IP market is in dire need for a more competitive and secure IP licensing scheme to flourish. In this paper, a pragmatic security protocol that could support the licensing of IP cores on a per-device basis without the need for a contractual agreement of an external trusted third party, large bandwidth, and complicated flow of communications is proposed. Besides assuring that the incentives of all stakeholders involved in the FPGA IP core trading are not scarified if not enhanced, the proposed protocol guarantees total secrecy and integrity of licensed IP cores based on available primitives of contemporary SRAM-based FPGA devices. It also prohibits the implementation of protected IP cores on unscreened excess devices and counterfeit chips sold in the gray market. The adoption of IP fingerprinting further deters the IP licensees from abusing the IP cores. Using the self-reconfiguration feature of FPGAs today, the resources consumed by the control module used for the implementation of protected IP cores on the authorized device are temporary and marginal. Chip-Hong Chang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Exploiting Process Variations and Programming Sensitivity of Phase Change Memory for Reconfigurable Physical Unclonable FunctionsabstractPhysical unclonable function (PUF) leverages the immensely complex and irreproducible nature of physical structures to achieve device authentication and secret information storage. To enhance the security and robustness of conventional PUFs, reconfigurable physical unclonable functions (RPUFs) with dynamically refreshable challenge-response pairs (CRPs) have emerged recently. In this paper, we propose two novel physically reconfigurable PUF (P-RPUF) schemes that exploit the process parameter variability and programming sensitivity of phase change memory (PCM) for CRP reconfiguration and evaluation. The first proposed PCM-based P-RPUF scheme extracts its CRPs from the measurable differences of the PCM cell resistances programmed by randomly varying pulses. An imprecisely controlled regulator is used to protect the privacy of the CRP in case the configuration state of the RPUF is divulged. The second proposed PCM-based RPUF scheme produces the random response by counting the number of programming pulses required to make the cell resistance converge to a predetermined target value. The merging of CRP reconfiguration and evaluation overcomes the inherent vulnerability of P-RPUF devices to malicious prediction attacks by limiting the number of accessible CRPs between two consecutive reconfigurations to only one. Both schemes were experimentally evaluated on 180-nm PCM chips. The obtained results demonstrated their quality for refreshable key generation when appropriate fuzzy extractor algorithms are incorporated. Le Zhang 0001, Zhi-Hui Kong, Chip-Hong Chang, Alessandro Cabrini, Guido Torelli |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2013 | Thermal simulator of 3D-IC with modeling of anisotropic TSV conductance and microchannel entrance effectsabstractThis paper presents a fast and accurate steady state thermal simulator for heatsink and microfluid-cooled 3D-ICs. This model considers the thermal effect of TSVs at fine-granularity by calculating the anisotropic equivalent thermal conductances of a solid grid cell if TSVs are inserted. Entrance effect of microchannels is also investigated for accurate modeling of microfluidic cooling. The proposed thermal simulator is verified against commercial multiphysics solver COMSOL and compared with Hotspot and 3D-ICE. Simulation results shows that for heatsink cooling, the proposed simulator is as accurate as Hotspot but runs much faster at moderate granularity. For microfluidic cooling, our proposed simulator is much more accurate than 3D-ICE in its estimation of steady state temperature and thermal distribution. Hanhua Qian, Hao Liang 0003, Chip-Hong Chang, Wei Zhang 0012, Hao Yu 0001 |
ASP-DAC | 3 |
| 2013 | Cluster-based distributed active current timer for hardware Trojan detectionabstractWith the globalization of integrated circuit (IC) design and fabrication, there is a growing concern on the devastating impact of subverted chip supply. This paper presents a current sensing circuit that converts the current activity on local power grid to a timing pulse to detect if an IC is Trojan-infected. This new approach increases the Trojan detection sensitivity by combining the switching activity and path sensitization abnormalities into a single side-channel signal that can be easily monitored by existing scan test structure. One main advantage of the proposed regional Trojan detector is that the current comparator threshold can be calibrated against the quiescent current noise floor to reduce the impacts of process variations. Experiments are performed on a Trojan-infected benchmark circuit to demonstrate the feasibility of the proposed technique. Yuan Cao 0003, Chip-Hong Chang, Shoushun Chen |
ISCAS | 2 |
| 2013 | A signed integer programmable power-of-two scaler for {2n-1, 2n, 2n+1} RNSabstractScaling is used for wordlength reduction in many DSP algorithms. Variable scaler provides greater flexibility than fixed scaler and allows for more efficient utilization of the dynamic range. Programmable magnitude scaling is extremely difficult to perform in RNS, let alone scaling of signed integer. To the best of our knowledge, no programmable signed integer scaler in RNS has been reported. What make sign integer scaling challenging is that sign detection is itself a problematic operation in RNS. This paper presents an algorithm that can selectively scale a signed integer in the three moduli set RNS {2n-1, 2n, 2n+1} by a power-of-two factor. The proposed architecture can be deemed as a signed RNS extension of the unsigned integer programmable power-of-two scaler for {2n-1, 2n, 2n+1}. On average, it reduces the area and total power consumption of the unsigned version by 27.4% and 30.3% with an average delay extension of slightly less than 20% to adjust the range of a residue after magnitude scaling. Jeremy Yung Shern Low, Thian Fatt Tay, Chip-Hong Chang |
ISCAS | 3 |
| 2013 | Microchannel splitting and scaling for thermal balancing of liquid-cooled 3DICabstractThis paper introduces the notion of channel splitting to augment the scaling of microchannels for a balanced microflu-idic cooling of 3DIC. The idea is to place appropriate number of channel splitters along various sections of microchannels to reduce the convective resistances at potential hotspots. The increases in pressure drops due to channel splitting are then redistributed by scaling the channel widths to match the coolant flowrates with the power distribution, which is usually nonuniform in practice. Unlike the existing techniques, this thermal balancing method requires no extra pump or valve. Only the customization of etching masks is needed for the deposition of silicon splitters and different sizes of microchannels. Experiment on a 4-layer multicore 3DIC stack shows that the proposed microchannel design technique can effectively reduce both the maximum temperature and thermal gradient in the 3D circuit. Hanhua Qian, Chip-Hong Chang |
ISCAS | 2 |
| 2013 | PCKGen: A Phase Change Memory based cryptographic key generatorabstractPhysical Unclonable Function (PUF) is widely known as an effective countermeasure to withstand non-invasive computational attacks as well as invasive tempering attacks on trusted computing systems. However, vast majority of the PUFs reported to-date are defined by static Challenge-Response Pairs (CRPs) with inferior security. In this paper, we propose a novel design of dynamically reconfigurable PUF based on Phase Change Memory (PCM) technology to yield refreshed cryptographic keys whenever the need arises to achieve enhanced security. A dedicated circuit framework is also introduced to reinforce the diversity of the CRP sets and improve the stability of the proposed PUF. Extensive simulation results show that our proposed work promises a clean delineation from the security bottlenecks faced by the state-of-the-art PUF designs. Le Zhang 0001, Zhi-Hui Kong, Chip-Hong Chang |
ISCAS | 3 |
| 2013 | An efficient channel clustering and flow rate allocation algorithm for non-uniform microfluidic cooling of 3D integrated circuits
Hanhua Qian, Chip-Hong Chang, Hao Yu 0001 |
Integr. | 2 |
| 2013 | Efficient VLSI Implementation of $2^{{n}}$ Scaling of Signed Integer in RNS ${\{2^{n}-1, 2^{n}, 2^{n}+1\}}$abstractScaling is a problematic operation in residue number system (RNS) but a necessary evil in implementing many digital signal processing algorithms for which RNS is particularly good. Existing signed integer RNS scalers entail a dedicated sign detection circuit, which is as complex as the magnitude scaling operation preceding it. In order to correct the incorrectly scaled negative integer in residue form, substantial hardware overheads have been incurred to detect the range of the residues upon magnitude scaling. In this brief, a fast and area efficient 2nsigned integer RNS scaler for the moduli set {2n-1, 2n, 2n+1} is proposed. A complex sign detection circuit has been obviated and replaced by simple logic manipulation of some bit-level information of intermediate magnitude scaling results. Compared with the latest signed integer RNS scalers of comparable dynamic ranges, the proposed architecture achieves at least 21.6% of area saving, 28.8% of speedup, and 32.5% of total power reduction for n ranging from 5 to 8. Thian Fatt Tay, Chip-Hong Chang, Jeremy Yung Shern Low |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Pipelined adder graph optimization for high speed multiple constant multiplicationabstractThis paper addresses the direct optimization of pipelined adder graphs (PAGs) for high speed multiple constant multiplication (MCM). The optimization opportunities are described and a definition of the pipelined multiple constant multiplication (PMCM) problem is given. It is shown that the PMCM problem is a generalization of the MCM problem with limited adder depth (AD). A novel algorithm to solve the PMCM problem heuristically, called RPAG, is presented. RPAG outperforms previous methods which are based on pipelining the solutions of conventional MCM algorithms. A flexible cost evaluation is used which enables the optimization for FPGA or ASIC targets on high or low abstraction levels. Results for both technologies are given and compared with the most recent methods. Even for the special case of limited AD it is shown that RPAG often produces better results compared to the prominent Hcubalgorithm with minimal total AD constraint. Martin Kumm, Peter Zipf, Mathias Faust, Chip-Hong Chang |
ISCAS | 4 |
| 2012 | A fast and compact circuit for integer square root computation based on Mitchell logarithmic methodabstractA novel non-iterative circuit for computing integer square root based on logarithm is proposed in the paper. Mitchell's methods are used for the logarithmic and antilogarithmic conversions. The proposed method merges two conversion stages into a single one to achieve better accuracy with a compact architecture. Hence, the circuit size and latency are reduced. Compared to an existing design based on the modified Dijkstra algorithm used in a coherent receiver, the proposed design is either 8 times smaller or 9 times faster for 16-bit integer input. Joshua Yung Lih Low, Ching-Chuen Jong, Jeremy Yung Shern Low, Thian Fatt Tay, Chip-Hong Chang |
ISCAS | 5 |
| 2012 | State encoding watermarking for field authentication of sequential circuit intellectual propertyabstractThis paper proposes a new watermarking scheme for intellectual property (IP) protection of sequential circuits. The method embeds the watermark by encoding the state variables as opposed to modifying the states and edges of state-transition graph (STG) in conventional finite state machine (FSM) watermarking schemes. It has the merits of being applicable to gate-level implementation of sequential circuit without the need to extract the STG; the authorship can be easily authenticated in the field as the watermark is embedded based on testability directed partitioning; and the watermark is hard to be erased by test structure removal and conceivable re-synthesis attacks as it is globally embedded before synthesis and optimization. The scheme has low probability of coincidence and very low logic overhead as evinced by the experimental results on ISCAS benchmark circuits and comparison with the most relevant state-of-the-art IP watermarking scheme. Chip-Hong Chang |
ISCAS | 2 |
| 2011 | A hybrid watermarking scheme for sequential functionsabstractMost watermarking schemes for intellectual property (IP) protection embed authorship information at a single design abstraction level. Effective means to directly verify the watermark distributed at the downstream designs are lacking, particularly after the IP core is packaged into chip. This paper proposes a hybrid scheme for watermarking sequential designs. At behavioral level, the finite state machine (FSM) is watermarked by interweaving the watermark bits into the outputs of its existing and unspecified transitions. During the Design-for-Testability (DfT) process, the scan chain that incorporates the watermarked FSM as one of the test kernels is further watermarked to provide a second level of protection. It enables the watermark on the FSM to be publicly detectable by the legitimate end users off chip via the scan test. Experimental results on ISCAS'89 and LGSynth'93 benchmark circuits show that the watermarked circuits have acceptably low overheads with stronger authorship proof. Aijiao Cui, Chip-Hong Chang |
ISCAS | 2 |
| 2011 | Bit-parallel Multiple Constant Multiplication using Look-Up Tables on FPGAabstractThe research on optimization of Multiple Constant Multiplication (MCM) during the last two decades has been focusing mainly on common subexpression elimination and reduced adder graph algorithms when bit-parallel computation is required. The advancement of FPGA technology enables the implementation of complex MCM instances on FPGA, but the shift-and-add network implementation does not make full use of the fundamental resources of FPGA, like the Look-Up Tables (LUT). Since bit-serial implementation optimized for FPGA is slow, an attempt for bit-parallel LUT-based implementation for single constant multiplication has been made. This paper extends this LUT-based method to multiple constant multiplications. It presents an interesting insight and unexpected outcome that the maximal number of LUTs required can be limited far below the theoretical number by mere enumeration without considering the legitimacy of all possible output combinations. Simulation results show that the required logic slices are comparable to the traditional adder-based MCM optimization methods while the delay is reduced by approximately 33%. The advantages are more prominent with increasing number of constants and the bit width used for their representation. Mathias Faust, Chip-Hong Chang |
ISCAS | 2 |
| 2011 | A new RNS scaler for {2n - 1, 2n, 2n + 1}abstractThis paper presents an efficient RNS scaling algorithm for the balanced special moduli set {2n-1, 2n, 2n+1}. By exploiting the relationship between the scaling constant and the residues of the three-moduli set using the New Chinese Remainder Theorem I (New CRT-I), the complicated modulo reduction operations for large integer scaling in RNS can be greatly simplified. The scaling constant has been chosen as 2n(2n+1)such that all residues of the scaled integer are identical and equal to the scaled integer output. This is particularly useful as no expensive and slow residue-to-binary converter is required for interfacing with conventional number system after the digital signal processing and scaling in RNS domain. The scaling error occurs only conditionally and is proven to be at most unity. The proposed design can be implemented entirely based on full adders with complexity commensurate with a multi-operand modulo 2n-1 adder. Its area-time complexity is at least 86% lower than one of the fastest ROM-based scaler designs for the same moduli set over a wide dynamic range of 15 bits and above. Jeremy Yung Shern Low, Chip-Hong Chang |
ISCAS | 2 |
| 2011 | A simple radix-4 Booth encoded modulo 2n+1 multiplierabstractAn area-efficient diminished-1 modulo 2n+1 multiplier with radix-4 modified Booth encoding is proposed. The proposed approach minimizes the number of Booth encoder and Booth decoder blocks required for partial product generation. Its correction factor is decomposed into a multiplier- dependent dynamic bias and a multiplier-independent static bias. The dynamic bias can be generated by hardwiring the outputs of the Booth encoder to appropriate bit positions, while the sum of the static bias and other multiplier-independent bias has been reduced to a simple binary word of alternate ones and zeros. For n = 40, the proposed multiplier achieves an area saving of 27% and a power reduction of 18.5% over the non- encoded modulo 2n+l multiplier. The proposed multiplier exhibits an area saving and an average power reduction of 52% and 62% respectively over the existing Booth encoded modulo 2n+l multiplier for the same n. The energy-delay product analysis indicates that the proposed multiplier provides an optimized trade-off between power consumption and delay. Ramya Muralidharan, Chip-Hong Chang |
ISCAS | 2 |
| 2011 | A Robust FSM Watermarking Scheme for IP Protection of Sequential Circuit DesignabstractFinite state machines (FSMs) are the backbone of sequential circuit design. In this paper, a new FSM watermarking scheme is proposed by making the authorship information a non-redundant property of the FSM. To overcome the vulnerability to state removal attack and minimize the design overhead, the watermark bits are seamlessly interwoven into the outputs of the existing and free transitions of state transition graph (STG). Unlike other transition-based STG watermarking, pseudo input variables have been reduced and made functionally indiscernible by the notion of reserved free literal. The assignment of reserved literals is exploited to minimize the overhead of watermarking and make the watermarked FSM fallible upon removal of any pseudo input variable. A direct and convenient detection scheme is also proposed to allow the watermark on the FSM to be publicly detectable. Experimental results on the watermarked circuits from the ISCAS'89 and IWLS'93 benchmark sets show lower or acceptably low overheads with higher tamper resilience and stronger authorship proof in comparison with related watermarking schemes for sequential functions. Aijiao Cui, Chip-Hong Chang, Sofiène Tahar, Amr Talaat Abdel-Hamid |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | Bayesian Separation With Sparsity Promotion in Perceptual Wavelet Domain for Speech Enhancement and Hybrid Speech RecognitionabstractSpeech recognition accuracy can be improved by the removal of noise. However, errors in the estimated signal components can also obscure the recognition. This paper presents a framework of wavelet-based techniques to harness the automatic speech recognition performance in the presence of background noise. The proposed robust speech recognition system is realized by implementing speech enhancement preprocessing, feature extraction, and a hybrid speech recognizer in the time–frequency space. A perceptual wavelet filterbank using a fixed base to imitate the human perceptual modus of speech is developed to capture the most discriminative information in the time–frequency plane. To minimize the mismatch between the training and testing conditions of the classifier, a Bayesian scheme is applied in a wavelet domain to separate the speech and noise components in the proposed iterative speech enhancement algorithm. The nonphonetic information is discarded while the more critical speech features are extracted and represented by the wavelet coefficients. The denoised wavelet features are fed to the hybrid classifier founded on a hidden Markov model (HMM). The intrinsic limitation of the HMM is overcome by augmenting it with a wavelet support vector machine. This hybrid and hierarchical design paradigm improves the recognition performance by combining the advantages of different methods into an integral system. The continuous digit speech recognition experiments conducted with the proposed framework show promising results. It significantly improves the recognition performance at a low signal-to-noise ratio (SNR) without causing a poorer performance at a high SNR. Chip-Hong Chang |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2011 | A High Bit Rate Serial-Serial Multiplier With On-the-Fly Accumulation by Asynchronous CountersabstractA novel approach of designing serial-serial hybrid multiplier is proposed for applications with high data sampling rate ( ≥4 GHz). The conventional way of partial product formation is revamped. Our proposed technique effectively forms the entire partial product matrix in justnsampling cycles for ann×nmultiplication instead of at least 2ncycles in the conventional serial-serial multipliers. It achieves a high bit sampling rate by replacing conventional full adders and 5:3 counters with asynchronous 1's counters so that the critical path is limited to only an and gate and a D flip-flop (DFF). The use of 1's counter to column compress the partial products preliminarily reduces the height of the partial product matrix fromnto [log2n] +1, resulting in a significant complexity reduction of the resultant adder tree. The proposed hybrid column compressed multiplier consists of a serial-serial data accumulation unit and a parallel carry save adder (CSA) array that occupies approximately 35% and 58% less silicon area than the full CSA array multiplier with operands of wordlength 32 × 32 and 64 × 64, respectively. The post-layout simulation results based on 90-nm seven metal single poly CMOS process technology shows that our 64 × 64 multiplier dissipates 39% less average power at a sampling rate of 4 GHz, and has only 11% additional delay penalty to complete a multiplication compared to the conventional fully parallel CSA array multiplier. Manas Ranjan Meher, Ching-Chuen Jong, Chip-Hong Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Minimal Logic Depth adder tree optimization for Multiple Constant MultiplicationabstractResearch on optimization of fixed coefficient FIR filters modeled as Multiple Constant Multiplication (MCM) has been ongoing for two decades. An analysis of Minimal Signed Digit (MSD) reveals that potential good solutions are omitted by Common Subexpression Elimination (CSE) algorithms as they are hidden in the MSD representations. Some CSE algorithms ensure that all coefficients are implemented at minimal Logic Depth (LD) which is advantageous from power saving perspective. Imposing this requirement on a graph dependant (GD) algorithm reduces the search space as well as the runtime. It also eliminates the long critical path of GD algorithm. This paper presents a minimal logic depth GD algorithm which requires no lookup table. Simulation results show that it has lower number of adders than CSE algorithms while having the minimal logic depth. For all filters tested, it consumes less switching power than the latest LD constrained GD methods based on the Glitch Path Count and Glitch Path Score metrics. Mathias Faust, Chip-Hong Chang |
ISCAS | 2 |
| 2010 | A novel counter-based low complexity inner-product architecture for high speed inputsabstractThis paper presents a new methodology of multiplierless implementation of inner-product computation. The inner-product computation is decomposed to form an architecture that facilitates an efficient serial accumulation of the 1's in the partial product matrix of each multiplication of a pair of elements from the input vectors. The 1's that appear at each partial product position are accumulated by a serial D flip flop (DFF) based 1's counter. This accumulation stage reduces the column height of the partial product matrix by transforming L vertical bits to ⌊log2L⌋ + 1 horizontal bits at each coordinate of the partial product matrix. The counter outputs are further summed using carry save addition based on Dadda's reduction algorithm followed by final carry propagating addition to obtain the final inner-product result. As the counter can operate at a frequency of 2 GHz when implemented on TSMC 0.18 μm CMOS process, the accumulation time is reduced significantly. With simpler partial product reduction tree, the hardware complexity is substantially reduced while the throughput is maintained to be comparable or better than many existing parallel inner-product computation architectures. The synthesis results of the proposed inner-product architecture for various inner-product lengths show that it has lower area-delay-product than many existing architectures. Manas Ranjan Meher, Ching-Chuen Jong, Chip-Hong Chang, Jeremy Yung Shern Low |
ISCAS | 3 |
| 2010 | Fast hard multiple generators for radix-8 Booth encoded modulo 2n-1 and modulo 2n+1 multipliersabstractHard multiple generation is the bottleneck operation in radix-8 Booth encoded modulo 2n- 1 and modulo 2n+ 1 multipliers. In this paper, fast hard multiple generators for the moduli 2n- 1 and 2n+ 1 are proposed. They are implemented as parallel-prefix structures based on the simplified carry equations. Synthesis results based on TSMC 0.18μm, 1.8V CMOS standard-cell library show that the proposed modulo 2n- 1 hard multiple generator reduces the critical path delay of the fastest general-purpose modulo 2n- 1 adder by 12% and 10% for n = 8 and n = 64, respectively. Compared to the smallest modulo 2n- 1 adder, the proposed design leads to 19% and 12% savings in silicon area for n = 8 and n = 64, respectively. The proposed modulo 2n+ 1 hard multiple generator also has the least critical path delay among the existing modulo 2n+ 1 adders. Ramya Muralidharan, Chip-Hong Chang |
ISCAS | 2 |
| 2010 | On "A New Common Subexpression Elimination Algorithm for Realizing Low-Complexity Higher Order Digital Filters"abstractA thorough analysis of the paper above revealed several controversial arguments about the superiority of binary representation over canonical signed digits (CSD) for common subexpression elimination (CSE). It was improper to model the number of logic operators (LO) required after CSE as a linear sum of independently weighted numbers of nonzero bits, common subexpressions and unpaired bits. The logic depth (LD) penalty of binary CSE had been deemphasized by the errors in the reported LD. This comment corrects the LD of contention resolution algorithm, and points out some contradictions with reference to the latest experimentation of binary, CSD and minimal signed digit number representations for CSE. Upon correcting the error in the reported filter lengths for different stopband attenuations of digital advanced mobile phone system specification, the LO and LD data of the CSE algorithms compared in the above paper are recalculated using the corrected filter coefficient sets. Chip-Hong Chang, Mathias Faust |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2010 | A Low Error and High Performance Multiplexer-Based Truncated MultiplierabstractThis paper proposes a novel adaptive pseudo-carry compensation truncation (PCT) scheme, which is derived for the multiplexer based array multiplier. The proposed method yields low average error among existing truncation methods. The new PCT based truncated array multiplier outperforms other existing truncated array multipliers by as much as 25% in terms of silicon area and delay, and consumes about 40% less dynamic power than the full-width multiplier for 32-bit operation. The proposed truncation scheme is applied to an image compression algorithm. Due to its low truncation error, the mean square errors (MSE) of various reconstructed images are found to be comparable to those obtained with full-precision multiplication. Chip-Hong Chang, Ravi Kumar Satzoda |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Time-multiplexed Data Flow Graph for the Design of Configurable Multiplier BlockabstractThis paper proposes a new design methodology to reduce the logic complexity of reconfigurable multiplier block (ReMB). The minimization problem is modeled as a scheduled time-multiplexed data flow graph (TDFG). To reduce the number of operators to be scheduled in the DFG, the most dominant common subexpressions are greedily identified and eliminated based on the subexpressions' frequencies which are updated dynamically in the optimization process. High level synthesis algorithm is then employed to perform the scheduling of operators to control steps. By binding the compatible operators in the same control steps, more operators can be saved. Two design examples are used to demonstrate the effectiveness of the proposed algorithm. On average, the logic complexity of the proposed ReMB design is about 19% lower than that of the classical ReMB methods, and 7% lower than that of the latest and most competitive ReMB design methodology. Jiajia Chen 0002, Chip-Hong Chang, Ching-Chuen Jong |
ISCAS | 2 |
| 2009 | New Power Index Model for Switching Power Analysis from Adder Graph of FIR filterabstractEfficient power modeling of a generic class of digital circuits is crucial to the analysis and development of optimization algorithms of power efficient design. This paper proposes a new switching power index model for the power analysis of multiplier-block based FIR filter. Unlike the existing glitch path count (GPC) and glitch path score (GPS), the proposed power index (PI) model takes into account correlated input switching activity propagations of full adders within and across adders of different widths, as well as the variation of load capacitances due to sharing of adders in reduced adder graph. Dynamic power simulations of several benchmark filters in an ASIC design flow show that this PI measure is more closely correlated with the actual dynamic power dissipation than the existing GPC and GPS models. Jiajia Chen 0002, Chip-Hong Chang, Hanhua Qian |
ISCAS | 2 |
| 2009 | An Improved Publicly Detectable Watermarking Scheme based on Scan Chain OrderingabstractThis paper proposes an improved version of watermarking scheme at the Design-for-testability (DfT) stage for VLSI Intellectual Property (IP) protection. The improved scheme overcomes the weaknesses of previous scan chain watermarking scheme by imposing the extra ordering constraints generated by the IP owner's signature on all scan flip-flops impartially. IP authorship can be publicly authenticated in the field by injecting a given test vector and matching a permuted output response vector against a transformed reference pattern. Both the output response and the reference sequence are related to a pseudorandom sequence generated by a public-key cryptographic algorithm. Experimental results show that the improved method has a low probability of coincidence and low test power overhead. Aijiao Cui, Chip-Hong Chang |
ISCAS | 2 |
| 2009 | Optimization of Structural Adders in Fixed Coefficient Transposed Direct Form FIR FiltersabstractOver the last two decades, fixed coefficient FIR filters were generally optimized by minimizing the number of adders required to implement the multiplier block in the transposed direct form filter structure. In this paper, an optimization method for the structural adders in the transposed tapped delay line is proposed. Although additional registers are required, an optimal trade-off can be made such that the overall combinational logic is reduced. For a majority of taps, the delay through the structural adder is shortened except for the last tap. The one full adder delay increase for the last optimized tap is tolerable as it does not fall in the critical path in most cases. The criterion for which area reduction is possible is analytically derived and an area reduction of up to 4.5% for the structural adder block of three benchmark filters is estimated theoretically. The saving is more prominent as the number of taps grows. Actual synthesis results obtained by synopsys design compiler with 0.18 mum TSMC CMOS libraries show a total area reduction of up to 13.13% when combined with common subexpression elimination. In all examples, up to 11.96% of the total area saved were due to the reduction of structural adder costs by our proposed method. Mathias Faust, Chip-Hong Chang |
ISCAS | 2 |
| 2009 | A Compact Current Mode Neuron Circuit with Gaussian Taper Learning CapabilityabstractIn this paper, an analog current mode implementation of a neuron circuit capable of performing real Gaussian neighborhood taper learning is presented. The neuron cell is compacted with a reusable multiplier that can function as squarer and multiplier for Euclidean and topological distances calculation as well as for Gaussian function characteristics with adjustable learning rate. A four-neuron self-organizing map (SOM) with three dimensional input data is designed and simulated using CSM 0.18 mum technology to demonstrate the learning control and neighborhood adaptation. The network can process 4.55 million vectors per second with a minimum power consumption of 1.6 mW at 1.5 V. Chip-Hong Chang, Liter Siek |
ISCAS | 2 |
| 2009 | Fixed and Variable Multi-modulus Squarer Architectures for Triple Moduli base of RNSabstractThe performance of RNS relies heavily on efficient implementation of residue arithmetic units. In this paper efficient multi-modulus squarer architectures for the moduli 2n-1, 2nand 2n+1 are presented. Two variants of multi-modulus squarer architectures, i.e., fixed and variable multi-modulus architectures, are proposed. Synthesis results based on TSMC 0.18 mum CMOS standard cell implementation demonstrate the performance trade-off between the various designs. Compared to single-modulus architecture for n = 24, fixed multi-modulus architecture provides an area and power savings of 5% and 10%, respectively with similar delay. On the other hand, for the same n, variable multi-modulus architecture reduces the area and power dissipation by 50% and 18%, respectively at the expense of 15% increase in delay. Ramya Muralidharan, Chip-Hong Chang |
ISCAS | 2 |
| 2009 | High-Level Synthesis Algorithm for the Design of Reconfigurable Constant MultiplierabstractMultiplying a signal by a known constant is an essential operation in digital signal processing algorithms. In many application scenarios, an input or output signal is repeatedly multiplied by several predefined constants at different instances. These temporal redundancies can be exploited for the design of an efficient reconfigurable constant multiplier (RCM). An RCM achieves greater hardware savings than the conventional multiple constant multiplication architecture, limited only by the available latency of the subsystem. Motivated by a number of lucrative examples, this paper presents a new high-level design methodology for RCM. Common subexpressions in the preset constants represented in minimum signed-digit system are first eliminated to obtain a minimum depth multiroot directed acyclic graph (DAG). The DAG is converted into a primitive data flow graph (DFG) where mobile adders are identified. By scheduling each mobile adder into a control step within its legitimate time window with the minimum opportunity cost, mutually exclusive adders can be merged with significantly reduced adder and multiplexing cost. The opportunity cost for each scheduling decision is assessed by the probability displacement and disparity measures of the scheduled node as well as its predecessors and successors in the DFG. The algorithm is runtime efficient as exhaustive search for the best fusion of independently optimized constant multipliers has been avoided. Simulation results on randomly generated 12-b constant sets show that the solutions generated by the proposed algorithm are on average 19% to 25% more area-time efficient than the best reported solutions. Jiajia Chen 0002, Chip-Hong Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Intellectual property authentication by watermarking scan chain in design-for-testability flowabstractThis paper proposes an intellectual property (IP) protection scheme at the Design-for-Testability (DfT) stage of VLSI design flow. Additional constraints generated by the owner’s digital signature have been imposed on the NP-hard problem of ordering the scan cells to achieve a watermarked solution which minimizes the penalty on power and cost of testing. As only the order of the scan cells is varied, the number of test vectors for the desired fault coverage is not affected. The advantage of this scheme is the ownership legitimacy can be publicly authenticated on-site by IP buyers after the chip has been packaged by loading a specific verification code into the scan chain. We propose to integrate the scan chain watermarking with dynamic watermarking of the IP core to make the design hard-to-attack while the ownership is easy-to-trace. The proposed scheme is applied to an optimization instance of scan cell ordering targeting at test power reduction. The results on several MCNC benchmarks show that the watermarking scheme has a very low probability of solution coincidence and hence provides strong proof of authorship. Aijiao Cui, Chip-Hong Chang |
ISCAS | 2 |
| 2008 | Programmable LSB-first and MSB-first modular multipliers for ECC in GF(2m)abstractIn this paper, we propose programmable serialin parallel-out LSB-first and MSB-first modular multipliers for elliptic curve cryptosystems (ECCs). The proposed multipliers can operate in any arbitrary field GF(2m) such that m is less than a maximum field order M. A linear array of processing elements is designed with a parallel switching circuitry to incorporate programmability in the fixed order multipliers. The proposed architectures are qualitatively compared against existing programmable multipliers in terms of gate count, delay and latency. The application specific integrated circuit (ASIC) implementation of the proposed multipliers using TSMC 0.18μm standard cell library is also analyzed. Ravi Kumar Satzoda, Ramya Muralidharan, Chip-Hong Chang |
ISCAS | 3 |
| 2008 | IP Watermarking Using Incremental Technology Mapping at Logic Synthesis LevelabstractThis paper proposes an adaptive watermarking technique by modulating some closed cones in an originally optimized logic network (master design) for technology mapping. The headroom of each disjoint closed cone is evaluated based on its slack and slack sustainability. The notion of slack sustainability in conjunction with an embedding threshold enables closed cones in the critical path to be qualified as watermark hosts if their slacks can be better preserved upon remapping. The watermark is embedded by remapping only qualified disjoint closed cones randomly selected and templates constrained by the signature. This parametric formulation provides a means to capitalize on the headroom of a design to increase the signature length or strengthen the watermark resilience. With the master design, the watermarked design can be authenticated as in nonoblivious media watermarking. Experimental results show that the design can be efficiently marked by our method with low overhead. Aijiao Cui, Chip-Hong Chang, Sofiène Tahar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Watermarking for IP Protection through Template Substitution at Logic Synthesis LevelabstractFunctionality preservation and area-time overhead are two major concerns of VLSI designers when the watermarking technique is used to protect their intellectual property. For a given technology library and a logically synthesized circuit, replacing cells with the templates of the same function will not alter the topology and original function of the circuit but the performances of some datapaths may be affected. If the resultant circuit can still satisfy the synthesis constraints with an acceptably low overhead, watermarking through template substitution at logic synthesis level is feasible. In this paper, we propose an algorithm to select appropriate cells for the replacement to realize the watermarking scheme. The method is tested on a set of combinational MCNC benchmarks. The results show that this watermarking process can provide a sufficiently strong proof of authorship with trivial area overhead. Aijiao Cui, Chip-Hong Chang |
ISCAS | 2 |
| 2007 | Design of Low-Complexity FIR Filters Based on Signed-Powers-of-Two Coefficients With Reusable Common SubexpressionsabstractIn this paper, a new efficient algorithm is proposed for the synthesis of low-complexity finite-impulse response (FIR) filters with resource sharing. The original problem statement based on the minimization of signed-power-of-two (SPT) terms has been reformulated to account for the sharable adders. The minimization of common SPT (CSPT) terms that were considered in our proposed algorithm addresses the optimization of the reusability of adders for two major types of common subexpressions, together with the minimization of adders that are needed for the spare SPT terms. The coefficient set is synthesized in two stages. In the first stage, CSPT terms in the vicinity of the scaled and rounded canonical signed digit (CSD) coefficients are allocated to obtain a CSD coefficient set, with the total number of CSPT terms not exceeding the initial coefficient set. The balanced normalized peak ripple magnitude due to the quantization error is fulfilled in the second stage by a local search method. The algorithm uses a common-subexpression-based hamming weight pyramid to seek for low-cost candidate coefficients with preferential consideration of shared common subexpressions. Experimental results demonstrate that our algorithm is capable of synthesizing FIR filters with the least CSPT terms compared with existing filter synthesis algorithms. Chip-Hong Chang, Ching-Chuen Jong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | A Generalized Time-Frequency Subtraction Method for Robust Speech Enhancement Based on Wavelet Filter Banks Modeling of Human Auditory SystemabstractWe present a new speech enhancement scheme for a single-microphone system to meet the demand for quality noise reduction algorithms capable of operating at a very low signal-to-noise ratio. A psychoacoustic model is incorporated into the generalized perceptual wavelet denoising method to reduce the residual noise and improve the intelligibility of speech. The proposed method is a generalized time-frequency subtraction algorithm, which advantageously exploits the wavelet multirate signal representation to preserve the critical transient information. Simultaneous masking and temporal masking of the human auditory system are modeled by the perceptual wavelet packet transform via the frequency and temporal localization of speech components. The wavelet coefficients are used to calculate the Bark spreading energy and temporal spreading energy, from which a time-frequency masking threshold is deduced to adaptively adjust the subtraction parameters of the proposed method. An unvoiced speech enhancement algorithm is also integrated into the system to improve the intelligibility of speech. Through rigorous objective and subjective evaluations, it is shown that the proposed speech enhancement system is capable of reducing noise with little speech degradation in adverse noise environments and the overall performance is superior to several competitive methods. Chip-Hong Chang |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2006 | Maximum likelihood disjunctive decomposition to reduced multirooted DAG for FIR filter designabstractThis paper extols the virtues of information theoretic approach to the synthesis of reduced multirooted directed acyclic graph (DAG) representation for the multiplier block of FIR filters. The proposed maximum likelihood decomposition algorithm can be viewed as an efficient divide-and-conquer approach with dynamic tracking of the statistic of weight-two subexpressions. As isomorphic subgraphs of the resultant reduced multirooted binary partition tree (MBPT) represent common subexpressions, higher weight common subexpressions are eliminated implicitly in the graph synthesis process. Experimental results show that the proposed algorithm produce designs with good tradeoffs for low logic complexity and logic depth Chip-Hong Chang, Jiajia Chen 0002, A. Prasad Vinod 0001 |
ISCAS | 1 |
| 2006 | Stego-signature at logic synthesis level for digital design IP protectionabstractThis paper presents a logic re-synthesis method for embedding the IP designer information into a distributed copy of a master design that has been synthesized to meet the application constraints. Slack information of the master copy is used to identify seed cells and extract their kernels for watermark insertion at the logic synthesis level. The embedded watermark can be recovered by comparing the topological mismatches between the marked circuit and the master copy. We demonstrate the difficulty of embedding or removing the watermark. The method has been tested on several MCNC multi-level logic synthesis benchmarks. Experimental results show that the method possesses high embedding capacity with trivial quality overhead for the synthesized solution. Aijiao Cui, Chip-Hong Chang |
ISCAS | 2 |
| 2006 | A low-power, high-speed RB-to-NB converter for fast redundant binary multiplierabstractIn this paper, a power-delay efficient redundant binary (RB) to 2's complement number converter for RB multiplication is presented. A new conversion algorithm is proposed to fully exploit the redundancy of RB encoding for a VLSI efficient implementation. The hierarchical linear expansion of the carry equation creates a regular multi-level parallel structure which is well suited for implementation with a logarithmic depth hybrid carry-lookahead/carry-select (CLA/ CSL) adder. A special add-one circuit is also incorporated into the CSL circuit to further reduce its logic complexity. A 64-bit reverse converter is designed using TSMC 0.18 /spl mu/m CMOS Technology. Pre-layout HSPICE simulation of the proposed design shows that it is capable of completing a 64-bit conversion in 761 ps and dissipates merely 0.34 mW at a data rate of 100MHz and a supply voltage of 1.8V. Yajuan He, Chip-Hong Chang |
ISCAS | 2 |
| 2006 | A fast kernel for unifying GF(p) and GF(2m) Montgomery multiplications in a scalable pipelined architectureabstractModular multiplication in Galois Fields - GF(p) and GF(2/sup m/) is an ineluctable and time stumbling block in public key cryptosystems. Montgomery modular multiplication has emerged as a VLSI efficient implementation of this operation. In this paper, a new scalable and pipelined Montgomery multiplier architecture that unifies the two important finite fields, GF(p) and GF(2/sup m/), is presented. The proposed architecture has successfully reduced the slack of the Montgomery multiplication in GF(2/sup m/) without jeopardizing the timing of its operation in GF(p). Acceleration of multiplication in GF(2/sup m/) for all ranges of modulus and in GF(p) for higher precision modulus is made possible through a new dual field adder and processing unit which can be pipelined in a kernel. The proposed dual field adder has been optimized to operate in an existing architecture that has been retimed to overcome the conflicts for speeding up the pipelined architecture. The latency has been analytically formulated in terms of the input wordlength, modulus precision and number of pipeline stages to evaluate its total computation time. The processing unit has been implemented on FPGA and the experimental results show evidence of throughput rate and latency improvement over existing dual field processing unit. Ravi Kumar Satzoda, Chip-Hong Chang |
ISCAS | 2 |
| 2006 | A Kalman filter based on wavelet filter-bank and psychoacoustic modeling for speech enhancementabstractThis paper presents a subband adaptive filter based on wavelet filter-bank for speech enhancement. The adaptation of Kalman filter in wavelet domain has effectively reduced the non-stationary noise. A perceptual weighting filter exploiting the masking properties of psychoacoustic model is concatenated with the Kalman filter to further improve the intelligibility of speech. The proposed method owns its merits from the successful porting of Kalman filter into the wavelet domain so that speech analysis and enhancement can be carried out in time-frequency spectrum based on the human auditory model. Experimental results show that the speech enhancement system is capable of reducing noise with little speech degradation in adverse noise environments and the overall performance is superior to several competitive methods. Chip-Hong Chang |
ISCAS | 2 |
| 2006 | A novel hybrid neuro-wavelet system for robust speech recognitionabstractThis paper presents a new automatic speech recognition system featuring the application of wavelet transform to speech enhancement method based on multilayer perceptron (MLP) classifier with a hidden Markov model (HMM). With the features extracted from a wavelet packet transform, different speech utterances are effectively discriminated by local discriminant bases. The extracted features is further processed by a feed-forward subsystem, a discriminant function minimum based blind adaptive filter for noise cancellation, and an unvoiced speech enhancement. A MLP network is used as the classifier before the Viterbi recognizer. Simulation results in adverse environments showed that the proposed system can achieve the best independent word recognition rate of 96.21%. The recognition degraded gracefully when it was tested by deliberately contaminating the signal with noises from the NOISEX-92 database. Chip-Hong Chang |
ISCAS | 2 |
| 2006 | A generalized perceptual time-frequency subtraction method for speech enhancementabstractThis paper presents a new speech enhancement scheme to meet the strong demand for quality noise reduction at very low signal-to-noise ratios (SNR). The proposed method generalizes the spectral subtraction algorithm to correlate time-frequency domain information based on the auditory masking property. The psycho acoustic model is integrated with the unvoiced speech enhancement algorithm to improve the intelligibility of speech. The proposed perceptual wavelet transform has successfully resolved the frequency and temporal components of speech signal. Auditory masking of noise is modeled by thresholding wavelet transform coefficients with adaptive suppression of different types of noise and distortion. Rigorous performance evaluations show that the proposed system is capable of reducing noise with little speech degradation in adverse noise environments and the perceptual speech quality is superior than several competitive methods. Chip-Hong Chang |
ISCAS | 2 |
| 2006 | Improved differential coefficients-based low power FIR filters. Part I. FundamentalsabstractThis paper and its companion paper (entitled Part II - algorithm) together present techniques for low power realization of finite impulse response (FIR) filters using improved differential coefficients method (DCM). This paper presents the necessary foundation and terminology of the DCM. The companion paper describes our algorithm and presents design examples. In contrast to the conventional DCM that is formulated at the algorithm-level, our method is formulated at the architecture-level using dedicated shift-and-add-based coefficient multipliers in order to achieve considerable hardware reduction. By employing a differential coefficient-partitioning algorithm (DCPA), we show that the number of full adders and the net memory needed to implement the coefficient multipliers can be significantly reduced. The proposed method is combined with common subexpression elimination method for further reduction of complexity. Experimental results show the average reductions of full adder, memory and energy dissipated achieved by our method over the DCM are 40%, 35% and 50% respectively. A. Prasad Vinod 0001, Ankita Singla, Chip-Hong Chang |
ISCAS | 3 |
| 2006 | A new integrated approach to the design of low-complexity FIR filtersabstractOptimizing adder cost for the implementation of finite-impulse response (FIR) filter has been an area of active research. The prevailing algorithms for the design of fixed point FIR filters have decoupled the filter coefficient synthesis from common subexpression elimination. This paper presents a new algorithm for the design of low-complexity FIR filters with resource sharing to reduce the adder cost directly during the coefficient synthesis process. The original problem statement based on the minimization of signed-power-of-two (SPT) terms is recast to account for the physical adder cost oriented common signed-power-of-two (CSPT) terms. Experimental results demonstrate that our algorithm is capable of synthesizing FIR filters with the least CSPT terms compared with existing filter synthesis algorithms Chip-Hong Chang, Ching-Chuen Jong |
ISCAS | 2 |
| 2005 | Fuzzy-ART based adaptive digital watermarking scheme
Chip-Hong Chang, Zhi Ye, Mingyan Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | New adaptive color quantization method based on self-organizing mapsabstractColor quantization (CQ) is an image processing task popularly used to convert true color images to palletized images for limited color display devices. To minimize the contouring artifacts introduced by the reduction of colors, a new competitive learning (CL) based scheme called the frequency sensitive self-organizing maps (FS-SOMs) is proposed to optimize the color palette design for CQ. FS-SOM harmonically blends the neighborhood adaptation of the well-known self-organizing maps (SOMs) with the neuron dependent frequency sensitive learning model, the global butterfly permutation sequence for input randomization, and the reinitialization of dead neurons to harness effective utilization of neurons. The net effect is an improvement in adaptation, a well-ordered color palette, and the alleviation of underutilization problem, which is the main cause of visually perceivable artifacts of CQ. Extensive simulations have been performed to analyze and compare the learning behavior and performance of FS-SOM against other vector quantization (VQ) algorithms. The results show that the proposed FS-SOM outperforms classical CL, Linde, Buzo, and Gray (LBG), and SOM algorithms. More importantly, FS-SOM achieves its superiority in reconstruction quality and topological ordering with a much greater robustness against variations in network parameters than the current art SOM algorithm for CQ. A most significant bit (MSB) biased encoding scheme is also introduced to reduce the number of parallel processing units. By mapping the pixel values as sign-magnitude numbers and biasing the magnitudes according to their sign bits, eight lattice points in the color space are condensed into one common point density function. Consequently, the same processing element can be used to map several color clusters and the entire FS-SOM network can be substantially scaled down without severely scarifying the quality of the displayed image. The drawback of this encoding scheme is the additional storage overhead, which can be cut down by leveraging on existing encoder in an overall lossy compression scheme. Chip-Hong Chang, Pengfei Xu 0009, Thambipillai Srikanthan |
IEEE Trans. Neural Networks | 1 |
| 2005 | Self-organizing topological tree for online vector quantization and data clusteringabstractThe self-organizing Maps (SOM) introduced by Kohonen implement two important operations: vector quantization (VQ) and a topology-preserving mapping. In this paper, an online self-organizing topological tree (SOTT) with faster learning is proposed. A new learning rule delivers the efficiency and topology preservation, which is superior of other structures of SOMs. The computational complexity of the proposed SOTT is O(log N) rather than O(N) as for the basic SOM. The experimental results demonstrate that the reconstruction performance of SOTT is comparable to the full-search SOM and its computation time is much shorter than the full-search SOM and other vector quantizers. In addition, SOTT delivers the hierarchical mapping of codevectors and the progressive transmission and decoding property, which are rarely supported by other vector quantizers at the same time. To circumvent the shortcomings of clustering performance of classical partition clustering algorithms, a hybrid clustering algorithm that fully exploit the online learning and multiresolution characteristics of SOTT is devised. A new linkage metric is proposed which can be updated online to accelerate the time consuming agglomerative hierarchical clustering stage. Besides the enhanced clustering performance, due to the online learning capability, the memory requirement of the proposed SOTT hybrid clustering algorithm is independent of the size of the data set, making it attractive for large database. Pengfei Xu 0009, Chip-Hong Chang, Andrew P. Paplinski |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2005 | A review of 0.18-μm full adder performances for tree structured arithmetic circuitsabstractThe general objective of our work is to investigate the area and power-delay performances of low-voltage full adder cells in different CMOS logic styles for the predominating tree structured arithmetic circuits. A new hybrid style full adder circuit is also presented. The sum and carry generation circuits of the proposed full adder are designed with hybrid logic styles. To operate at ultra-low supply voltage, the pass logic circuit that cogenerates the intermediate XOR and XNOR outputs has been improved to overcome the switching delay problem. As full adders are frequently employed in a tree structured configuration for high-performance arithmetic circuits, a cascaded simulation structure is introduced to evaluate the full adders in a realistic application environment. A systematic and elegant procedure to scale the transistor for minimal power-delay product is proposed. The circuits being studied are optimized for energy efficiency at 0.18-/spl mu/m CMOS process technology. With the proposed simulation environment, it is shown that some survival cells in stand alone operation at low voltage may fail when cascaded in a larger circuit, either due to the lack of drivability or unsatisfactory speed of operation. The proposed hybrid full adder exhibits not only the full swing logic and balanced outputs but also strong output drivability. The increase in the transistor count of its complementary CMOS output stage is compensated by its area efficient layout. Therefore, it remains one of the best contenders for designing large tree structured arithmetic circuits with reduced energy consumption while keeping the increase in area to a minimum. Chip-Hong Chang, Jiangmin Gu, Mingyan Zhang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | Efficient algorithms for common subexpression elimination in digital filter designabstractA contention resolution algorithm (CRA) is proposed for the common subexpression elimination of the multiplier block of the digital filter structure. CRA synthesizes common subexpressions of any Hamming weight to achieve an overall minimization with the emphasis that every logic depth increment must be accompanied by a reduction in logic complexity. A new data structure, called the admissibility graph is introduced to represent succinctly a set of coefficients; the admissible subexpressions are progressively labeled on the graph as either precedence or contention edges (or paths). The performance of CRA is evaluated based on benchmarked circuits and randomly generated coefficients. It is demonstrated that our algorithm outperforms several distinguished algorithms in both the logic depth and logic complexity. Chip-Hong Chang, Ching-Chuen Jong |
ICASSP (5) | 2 |
| 2003 | An adaptive initialization technique for color quantization by self organizing feature mapabstractAn unsupervised learning network, such as the self organizing feature map (SOFM), has been applied successfully to color classification for image compression and pattern recognition. Like other vector quantization algorithms, the reconstruction quality and adaptation rate of the SOFM are sensitive to the neuron initialization. We propose an efficient new initialization method, whereby an excess number of neurons is defined and the neurons are adaptively pruned, merged and split within their lattice according to the spatial distribution of the input color pixels. Comparisons with conventional gray scale initialization using subsampling and butterfly jumping sequences show that the proposed method obtains good initial code vectors that can accelerate the convergence of the SOFM and improve the reconstructed image quality significantly. Chip-Hong Chang, Thambipillai Srikanthan |
ICASSP (3) | 1 |
| 2003 | Low voltage, low power (5: 2) compressor cell for fast arithmetic circuitsabstractThis paper presents a new (5:2) compressor circuit capable of operating at ultra-low voltages. Its power efficacy is derived from the novel design of composite XOR-XNOR gate at transistor level. The new circuit eliminates the weak logic and threshold voltage drop problems, which are the main factors limiting the performance of pass transistor based circuits at low supply voltages. The proposed (5:2) compressor has been designed with special consideration on output drivability to ensure that it can function reliably at low voltages when these cells are employed in the tree structured multiplier and multiply-accumulator. Simulation results show that the proposed (5:2) compressor is able to function at supply voltage as low as 0.7 V, and. outperforms other (5:2) compressors constructed with various combinations of recently reported superior low-power logic cells. Jiangmin Gu, Chip-Hong Chang |
ICASSP (2) | 2 |
| 1997 | Forward and Inverse Transformations Between Haar Spectra and Ordered Binary Decision Diagrams of Boolean FunctionsabstractUnnormalized Haar spectra and Ordered Binary Decision Diagrams (OBDDs) are two standard representations of Boolean functions used in logic design. In this article, mutual relationships between those two representations have been derived. The method of calculating the Haar spectrum from OBDD has been presented. The decomposition of the Haar spectrum, in terms of the cofactors of Boolean functions, has been introduced. Based on the above decomposition, another method to synthesize OBDD directly from the Haar spectrum has been presented. Bogdan J. Falkowski, Chip-Hong Chang |
IEEE Trans. Computers | 2 |
| 1995 | Flexible optimization of fixed polarity Reed-Muller expansions for multiple and output completely and incompletely specified boolean functionsabstractNo abstract available. Chip-Hong Chang, Bogdan J. Falkowski |
ASP-DAC | 1 |
| 1995 | Generation of Multi-Polarity Arithmetic Transform from Reduced Representation of Boolean Functions
Bogdan J. Falkowski, Chip-Hong Chang |
ISCAS | 2 |
| 1994 | Efficient Algorithms for the Calculation of Arithmetic Spectrum from OBDD & Synthesis of OBDD from Arithmetic Spectrum for Incompletely Specified Boolean FunctionsabstractAn algorithm has been developed to calculate the arithmetic transform of Boolean functions from their Ordered Binary Decision Diagram (OBDD) representation. The method of decomposition of arithmetic spectral coefficients in terms of the cofactors of Boolean functions that resembles known Shannon decomposition of such functions has been introduced for the first time. Based on the above decomposition, a second new algorithm is presented to synthesize Ordered Binary Decision Diagrams directly from the arithmetic spectrum of Boolean functions.> Bogdan J. Falkowski, Chip-Hong Chang |
ISCAS | 2 |